← Research catalog

S04 / ADVERTISING AND MARKETING TECHNOLOGY

An uplift model still has to beat the existing policy.

At 20% capacity, the selected causal forest does not demonstrate a gain over response targeting on an independent advertising benchmark.

Evaluated public-data studyResearch code ↗Executive summary ↓

398,506 final records · 1,243 conversions · corrected Criteo v2.1 · profile-separated evaluation

01 / DECISION · EXECUTIVE SUMMARY

Require evidence before replacing response targeting.

A marketing leader wants to reach people whose behavior changes because of advertising, rather than simply people likely to buy. That is the right distinction to test. In this benchmark, however, the selected uplift model does not show a measurable advantage over a model that ranks likely converters.

9.78selected uplift policy
9.86response targeting
2.04expected random allocation

Estimated incremental benchmark conversions per 10,000 eligible records, each at 20% targeting capacity. These are estimates under benchmark assumptions, not observed caused-conversion counts.

What a leader should do next

Keep response targeting as a serious comparator. Before paying for a more complex system, test the change at the actual budget in a current randomized campaign with known costs. Evaluate independent outcomes, not the uplift model’s own predicted gains.

What this does not establish

The paired difference is -0.08 conversions per 10,000, with a 95% interval from -0.89 to +0.73. That is insufficient evidence of an uplift advantage; it is not proof of equality. Privacy subsampling prevents recovery of original advertiser incrementality or ROI.

02 / IMPLICATION

Choose the targeting system by its independent decision result.

The two targeted approaches have similar final estimates despite different modeling objectives. Their point estimates exceed expected random allocation, but the upgrade decision is the paired comparison with response targeting. The more complex method has not earned replacement on this test.

Capacity matters. The explorer preserves all saved budget points, including targeting nobody and everybody. The prespecified primary comparison remains 20%, so choosing a favorable point afterward does not turn a supplementary result into the main finding.

03 / EVIDENCE

Independent outcomes, matched budgets and visible uncertainty.

Final evaluation uses 398,506 records with 352,598 distinct anonymized feature profiles. There are 60,018 control records with 127 conversions, and 338,488 treatment-assigned records with 1,116 conversions. Rare control events limit precision.

ADVERTISING CAPACITY EXPLORER

Does a different ranking earn the same budget?

These are saved independent evaluation scores. Capacity changes the frozen ranking’s selection; it does not refit the model or observe a new experiment.

9.78benchmark conversions per 10,000
5.96 to 13.60conditional 95% interval
79,701selected evaluation records

Conversion estimates are incremental benchmark contrasts relative to no targeting, under the study’s identification assumptions. They are not observed counts of caused conversions or original advertiser effects.

Independent policy contrast across targeting capacitiesSelected policy and its conditional 95% band, response targeting and expected random targeting. Exact comparisons at the selected capacity follow.0.00%3.625%7.250%10.875%14.3100%Benchmark conversions per 10,000 eligible recordsTargeting capacity (%)
Copper and shaded band: selected policy. Navy: response targeting. Dashed slate: expected random allocation. Intervals are conditional and not simultaneous across the frontier. Scroll horizontally on narrow screens.
All policies · 20% capacity
PolicyContrast / 10,00095% interval
Doubly robust learner · L2 109.545.88 to 13.21
Doubly robust learner · L2 1009.666.02 to 13.30
Honest causal forest · leaf 1009.505.70 to 13.29
Honest causal forest · leaf 5009.785.96 to 13.60
Response-probability targeting9.866.06 to 13.67
Expected random allocation2.041.24 to 2.84

At 0%, every policy targets nobody and its contrast is zero. At 100%, all policies target everyone and share the same estimate. The primary model decision was fixed at 20%; the other capacity points are supplementary.

ASSUMED ECONOMICS / SCENARIO ONLY

Translate the benchmark into an explicit assumption.

Neither conversion value nor contact cost is observed. This calculation is a benchmark illustration, not original advertiser ROI or profit.

$777.93 illustrative net value per 10,000 eligible records

Conditional scenario range: $396.00 to $1,159.85. Negative values indicate assumed cost exceeds the estimated benchmark value at these inputs.

Calculation: benchmark conversion contrast × assumed value − selected fraction × 10,000 × assumed targeting cost. Fixed-price intervals exclude uncertainty in economics, privacy sampling, future populations and model refitting.

Does predicted uplift agree with independent outcomes?
Selected forest · predicted-effect deciles
DecileRecordsConversionsPredicted / 10,000Independent / 10,00095% interval
139,85115-4.060.63-4.46 to 5.73
239,8514-1.17-0.78-4.14 to 2.58
339,8512-0.740.67-0.16 to 1.50
439,8512-0.450.59-0.23 to 1.42
539,85110-0.111.15-2.48 to 4.79
639,851110.65-4.54-11.17 to 2.10
739,850151.364.542.29 to 6.79
839,850203.771.97-3.26 to 7.20
939,8508014.995.47-5.87 to 16.80
1039,8501084100.8892.3355.87 to 128.79

Groups are formed from predictions without using outcomes. Their wide intervals reflect sparse events; they are not an independent human judgment of individual treatment effects.

Sensitivity to adjustment and repeated profiles
Selected forest · frozen 20% policy
AnalysisSelected shareContrast / 10,00095% interval
Learned propensity, doubly robust20.0%9.785.96 to 13.60
Constant propensity, doubly robust20.0%10.206.54 to 13.86
Inverse propensity weighting20.0%10.085.91 to 14.26
One record per profile22.5%11.246.94 to 15.54

Keeping one row per matching profile changes the population and selected share to 22.5%. It preserves the original policy decisions; it is not a new same-capacity comparison. Matching profiles do not prove matching people.

Paired 95% profile-cluster normal intervals from independent held-out score contributions. Conditional on fitted models and rankings; unobserved campaign dependence and privacy-selection bias are not identified. The confidence bands do not resolve causal-identification uncertainty introduced by the release’s sampling.

04 / DATA

The corrected release, with its sampling limits intact.

Criteo’s corrected v2.1 file contains 13,979,592 records and twelve anonymized features, f0–f11. The older release had documented leakage and is excluded. This study keeps an outcome-independent sample of 1,995,142 records and places every matching profile in a single partition.

Model fitting uses 596,910 development records; validation uses 399,506; final evaluation uses 398,506. A further fixed development hash rule bounds computation. No raw data or individual predictions are redistributed.

Treatment means assigned advertising eligibility. Realized exposure, visit and conversion are not predictors. No dates, stable user identifiers, campaign strata or original monetary values are released. A chronological evaluation or original advertiser ROI cannot be inferred from these files.

Diemert, Betlei, Renaudin & Amini (2018), A Large Scale Benchmark for Uplift Modeling, AdKDD/TargetAd Workshop. Corrected publisher release and terms ↗

The source and derived aggregate research tables retain CC BY-NC-SA 4.0 attribution and no-warranty terms. Original analysis code is separate. The publisher does not endorse these findings.

05 / METHOD & LIMITATIONS

Cross-fitting for learning, independent scoring for the decision.

Three profile-separated development folds estimate outcome and treatment-propensity functions. A doubly robust learner predicts adjusted treatment contrasts; an honest causal forest uses separate observations for tree structure and leaf estimation. Two fixed settings per family are compared only on validation at 20% capacity.

The honest forest with minimum leaf size 500 wins validation. All four candidates remain in final reporting. An independent doubly robust score combines held-out outcomes with development-fitted nuisance functions; summing predicted uplift is never used as evidence of a policy’s success.

Final propensity scores range from 0.773 to 0.895, with none clipped. Assignment-prediction AUROC is 0.509. These are overlap/balance diagnostics, not proof that privacy sampling preserves exchangeability.

  • No dates, campaign IDs, demographic meanings or original experimental strata are released.
  • Rare conversion and control events can make capacity-specific contrasts imprecise.
  • Cluster intervals omit fitting uncertainty and are not simultaneous frontier guarantees.
  • Profile-hash sampling and bounded fitting do not establish full-source or current-market performance.
  • Exposure and visit are excluded as post-assignment variables.
  • No advertiser ROI, realized profit or fairness conclusion is identified.
  • Independent technical review is pending.

Frozen policy-evaluation protocol ↗

06 / CODE

Trace each budget point to an executed comparison.

Run S04-36e06a94-2716e1bf
Analysis commit 36e06a9402b561bcf051e186c11ae9aeabf8aa42

git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S04/study.py
uv run python scripts/report_s04.py
uv run python -W error -m unittest discover -s tests -v

Python 3.11.16 · EconML 0.17.0 · locked dependencies · fixed seeds. Checks cover score identities, independent policy outcomes, grouped folds, endpoint accounting and validation-only selection. Independent technical review is pending.