01 / DECISION · EXECUTIVE SUMMARY
Require evidence before replacing response targeting.
A marketing leader wants to reach people whose behavior changes because of advertising, rather than simply people likely to buy. That is the right distinction to test. In this benchmark, however, the selected uplift model does not show a measurable advantage over a model that ranks likely converters.
Estimated incremental benchmark conversions per 10,000 eligible records, each at 20% targeting capacity. These are estimates under benchmark assumptions, not observed caused-conversion counts.
What a leader should do next
Keep response targeting as a serious comparator. Before paying for a more complex system, test the change at the actual budget in a current randomized campaign with known costs. Evaluate independent outcomes, not the uplift model’s own predicted gains.
What this does not establish
The paired difference is -0.08 conversions per 10,000, with a 95% interval from -0.89 to +0.73. That is insufficient evidence of an uplift advantage; it is not proof of equality. Privacy subsampling prevents recovery of original advertiser incrementality or ROI.
02 / IMPLICATION
Choose the targeting system by its independent decision result.
The two targeted approaches have similar final estimates despite different modeling objectives. Their point estimates exceed expected random allocation, but the upgrade decision is the paired comparison with response targeting. The more complex method has not earned replacement on this test.
Capacity matters. The explorer preserves all saved budget points, including targeting nobody and everybody. The prespecified primary comparison remains 20%, so choosing a favorable point afterward does not turn a supplementary result into the main finding.
03 / EVIDENCE
Independent outcomes, matched budgets and visible uncertainty.
Final evaluation uses 398,506 records with 352,598 distinct anonymized feature profiles. There are 60,018 control records with 127 conversions, and 338,488 treatment-assigned records with 1,116 conversions. Rare control events limit precision.
ADVERTISING CAPACITY EXPLORER
Does a different ranking earn the same budget?
These are saved independent evaluation scores. Capacity changes the frozen ranking’s selection; it does not refit the model or observe a new experiment.
Conversion estimates are incremental benchmark contrasts relative to no targeting, under the study’s identification assumptions. They are not observed counts of caused conversions or original advertiser effects.
| Policy | Contrast / 10,000 | 95% interval |
|---|---|---|
| Doubly robust learner · L2 10 | 9.54 | 5.88 to 13.21 |
| Doubly robust learner · L2 100 | 9.66 | 6.02 to 13.30 |
| Honest causal forest · leaf 100 | 9.50 | 5.70 to 13.29 |
| Honest causal forest · leaf 500 | 9.78 | 5.96 to 13.60 |
| Response-probability targeting | 9.86 | 6.06 to 13.67 |
| Expected random allocation | 2.04 | 1.24 to 2.84 |
At 0%, every policy targets nobody and its contrast is zero. At 100%, all policies target everyone and share the same estimate. The primary model decision was fixed at 20%; the other capacity points are supplementary.
ASSUMED ECONOMICS / SCENARIO ONLY
Translate the benchmark into an explicit assumption.
Neither conversion value nor contact cost is observed. This calculation is a benchmark illustration, not original advertiser ROI or profit.
$777.93 illustrative net value per 10,000 eligible records
Conditional scenario range: $396.00 to $1,159.85. Negative values indicate assumed cost exceeds the estimated benchmark value at these inputs.
Calculation: benchmark conversion contrast × assumed value − selected fraction × 10,000 × assumed targeting cost. Fixed-price intervals exclude uncertainty in economics, privacy sampling, future populations and model refitting.
Does predicted uplift agree with independent outcomes?
| Decile | Records | Conversions | Predicted / 10,000 | Independent / 10,000 | 95% interval |
|---|---|---|---|---|---|
| 1 | 39,851 | 15 | -4.06 | 0.63 | -4.46 to 5.73 |
| 2 | 39,851 | 4 | -1.17 | -0.78 | -4.14 to 2.58 |
| 3 | 39,851 | 2 | -0.74 | 0.67 | -0.16 to 1.50 |
| 4 | 39,851 | 2 | -0.45 | 0.59 | -0.23 to 1.42 |
| 5 | 39,851 | 10 | -0.11 | 1.15 | -2.48 to 4.79 |
| 6 | 39,851 | 11 | 0.65 | -4.54 | -11.17 to 2.10 |
| 7 | 39,850 | 15 | 1.36 | 4.54 | 2.29 to 6.79 |
| 8 | 39,850 | 20 | 3.77 | 1.97 | -3.26 to 7.20 |
| 9 | 39,850 | 80 | 14.99 | 5.47 | -5.87 to 16.80 |
| 10 | 39,850 | 1084 | 100.88 | 92.33 | 55.87 to 128.79 |
Groups are formed from predictions without using outcomes. Their wide intervals reflect sparse events; they are not an independent human judgment of individual treatment effects.
Sensitivity to adjustment and repeated profiles
| Analysis | Selected share | Contrast / 10,000 | 95% interval |
|---|---|---|---|
| Learned propensity, doubly robust | 20.0% | 9.78 | 5.96 to 13.60 |
| Constant propensity, doubly robust | 20.0% | 10.20 | 6.54 to 13.86 |
| Inverse propensity weighting | 20.0% | 10.08 | 5.91 to 14.26 |
| One record per profile | 22.5% | 11.24 | 6.94 to 15.54 |
Keeping one row per matching profile changes the population and selected share to 22.5%. It preserves the original policy decisions; it is not a new same-capacity comparison. Matching profiles do not prove matching people.
Paired 95% profile-cluster normal intervals from independent held-out score contributions. Conditional on fitted models and rankings; unobserved campaign dependence and privacy-selection bias are not identified. The confidence bands do not resolve causal-identification uncertainty introduced by the release’s sampling.
04 / DATA
The corrected release, with its sampling limits intact.
Criteo’s corrected v2.1 file contains 13,979,592 records and twelve anonymized features, f0–f11. The older release had documented leakage and is excluded. This study keeps an outcome-independent sample of 1,995,142 records and places every matching profile in a single partition.
Model fitting uses 596,910 development records; validation uses 399,506; final evaluation uses 398,506. A further fixed development hash rule bounds computation. No raw data or individual predictions are redistributed.
Treatment means assigned advertising eligibility. Realized exposure, visit and conversion are not predictors. No dates, stable user identifiers, campaign strata or original monetary values are released. A chronological evaluation or original advertiser ROI cannot be inferred from these files.
Diemert, Betlei, Renaudin & Amini (2018), A Large Scale Benchmark for Uplift Modeling, AdKDD/TargetAd Workshop. Corrected publisher release and terms ↗
The source and derived aggregate research tables retain CC BY-NC-SA 4.0 attribution and no-warranty terms. Original analysis code is separate. The publisher does not endorse these findings.
05 / METHOD & LIMITATIONS
Cross-fitting for learning, independent scoring for the decision.
Three profile-separated development folds estimate outcome and treatment-propensity functions. A doubly robust learner predicts adjusted treatment contrasts; an honest causal forest uses separate observations for tree structure and leaf estimation. Two fixed settings per family are compared only on validation at 20% capacity.
The honest forest with minimum leaf size 500 wins validation. All four candidates remain in final reporting. An independent doubly robust score combines held-out outcomes with development-fitted nuisance functions; summing predicted uplift is never used as evidence of a policy’s success.
Final propensity scores range from 0.773 to 0.895, with none clipped. Assignment-prediction AUROC is 0.509. These are overlap/balance diagnostics, not proof that privacy sampling preserves exchangeability.
- No dates, campaign IDs, demographic meanings or original experimental strata are released.
- Rare conversion and control events can make capacity-specific contrasts imprecise.
- Cluster intervals omit fitting uncertainty and are not simultaneous frontier guarantees.
- Profile-hash sampling and bounded fitting do not establish full-source or current-market performance.
- Exposure and visit are excluded as post-assignment variables.
- No advertiser ROI, realized profit or fairness conclusion is identified.
- Independent technical review is pending.
06 / CODE
Trace each budget point to an executed comparison.
Run S04-36e06a94-2716e1bf
Analysis commit 36e06a9402b561bcf051e186c11ae9aeabf8aa42
git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S04/study.py
uv run python scripts/report_s04.py
uv run python -W error -m unittest discover -s tests -vPython 3.11.16 · EconML 0.17.0 · locked dependencies · fixed seeds. Checks cover score identities, independent policy outcomes, grouped folds, endpoint accounting and validation-only selection. Independent technical review is pending.