← Research catalog

S03 / RETAIL AND E-COMMERCE

When is a customer worth winning back?

The structured purchase model beats the flexible challenger overall. Seasonal errors still limit how confidently a campaign team can use it.

Evaluated public-data studyResearch code ↗Executive summary ↓

UK giftware retailer · December 2009–December 2011 · 5,155 final customers · three 90-day forecast windows

01 / DECISION & EXECUTIVE SUMMARY

Forecast who may return. Test separately whether contact changes that.

A marketing leader needs to understand future purchasing before committing a win-back budget. In this historical retailer, the BG/NBD purchase model is more accurate than the boosted challenger on the primary count-forecast score. It predicts that 37.9% of customer-windows will contain a purchase; 36.9% actually do. Boosting predicts 49.0%.

0.990 vs 1.072count deviance: structured model vs boosting · lower is better
14,38890-day customer-window forecasts · 5,155 distinct customers
10,569observed future purchase days

Keep the structured model as the reference.

The challenger increases primary forecast error by 0.082 deviance units (95% interval 0.068–0.096). This is a historical prediction comparison, not a percentage revenue gain. The simpler recent-purchase rule still has lower absolute count error, so no model wins every measure.

The pooled result hides seasonal misses.

BG/NBD overpredicts purchase days by 28.9% in the March window, then underpredicts by 19.5% in September. It estimates ordinary purchasing, not the additional purchases caused by a campaign. No randomized win-back intervention or observed permanent-churn label is available.

02 / IMPLICATION

A contact budget needs incremental value, not just likely buyers.

For the whole 90-day cohort, BG/NBD plus Gamma-Gamma predicts £380.54 gross positive spend per customer-window against £401.24 recorded. The model’s predicted purchase-day value is £490.83. At an assumed £1 contact cost and 30% contribution margin, break-even requires about 0.68 percentage points of additional purchase probability, assuming one extra purchase day.

That is a threshold for a future campaign experiment—not evidence that a campaign reaches it. This retailer includes wholesale buyers; gross purchase values are not generic consumer economics, net revenue or profit. Seasonal validation and a randomized contact test would be needed before claiming campaign value.

03 / EVIDENCE

Compare forecasting loss, calibration and uncertainty.

Three nonoverlapping 90-day windows begin March 1, June 1 and September 1, 2011. Customers recur across windows; 1,000 paired bootstrap draws keep each customer’s windows together. Intervals condition on the fitted models and these historical periods.

Final 90-day forecasts · 95% customer-bootstrap intervals
ModelCount deviance95% intervalPurchase BrierPredicted / observed spend
Recent 90-day purchasing7.4197.098–7.7320.2260.867
Historical RFM cohorts1.3281.266–1.3990.1871.485
BG/NBD + Gamma-Gamma0.9900.963–1.0180.1700.948
Hurdle count–spend boosting1.0721.051–1.0930.1901.227

Count deviance and Brier loss are lower when forecasts improve. The recent rule assigns zero expected purchases to customers who have been inactive for 90 days but sometimes return; log scoring floors those means at 1e-8. Its absolute count error is nevertheless 0.584 versus BG/NBD’s 0.621. Its revenue MAE is £315.81 versus £321.84 for BG/NBD.

BG/NBD’s predicted/observed purchase-count ratio is 1.055 (95% interval 1.029–1.082); boosting’s is 1.303 (1.271–1.336). Gross-spend ratios are 0.948 (0.885–1.009) and 1.227 (1.127–1.327). A well-calibrated pooled revenue total can still hide badly calibrated periods.

HISTORICAL CUSTOMER EXPLORER

Separate a likely purchase from an incremental purchase.

Select saved forecasts for a historical cohort. Cost and margin change only the hypothetical break-even calculation; they do not change evaluated accuracy or establish campaign lift.

14,388 customer-window observations · 5,155 distinct customers · 90-day horizon. Customers can recur across the three final forecast dates.

37.9%predicted chance of any purchase
36.9%observed windows with a purchase
0.78 / 0.73predicted / observed purchase days

Gross positive spend per customer-window: £380.54 predicted versus £401.24 observed. A purchase day combines all positive purchases on one date; gross value excludes returns and is not profit.

Latent model activity: 91.3%. This is different from purchasing within 90 days and is not observed retention. BG/NBD fixes activity at 100% for customers with no repeat purchase.

What would contact need to cause?

£1.00 ÷ (30% × £490.83) = 0.68 percentage points of incremental purchase probability.

Assumes one additional purchase day at the model’s predicted value. No win-back intervention was observed. This threshold is not expected lift or campaign ROI; ordinary predicted purchasing is not incremental response.

Nominal 90% predictive bands covered 97.8% of count outcomes (average width 2.41 purchase days) and 96.2% of spend outcomes (average width £1,237.62). These are empirical coverage rates for individual customer-windows, not confidence intervals for this cohort average.

Predicted and observed purchase-day distribution for the selected customer cohort012345+0%50%100%Future purchase days
Navy: model distribution from 2,000 predictive draws per observation. Copper: actual frequency. The last bin contains five or more purchase days; the distribution conditions on fitted population parameters.
90-day purchase distribution · selected cohort
Purchase daysPredictedObserved
062.1%63.1%
120.3%20.9%
28.9%8.3%
34.1%3.5%
42.0%1.8%
5+2.6%2.3%
90-day saved forecasts · All inactivity bands · All purchase histories
ModelPurchase probabilityExpected purchase daysGross spend
Recent 90-day purchasing27.1%0.684£348.04
Historical RFM cohorts44.7%0.976£595.72
BG/NBD + Gamma-Gamma37.9%0.775£380.54
Hurdle count–spend boosting49.0%0.957£492.17
Observed36.9%0.735£401.24

Historical cohorts are descriptive, with repeated customers and uneven seasonal performance. The primary comparison’s customer-bootstrap intervals appear above. These aggregate filters do not run training, produce individual contact lists or demonstrate causal effects.

Seasonal calibration and predictive interval coverage
Separate 90-day windows · a ratio of 1 matches observed totals
OriginModelCount ratioSpend ratio90% spend coverage
2011-03-01BG/NBD + Gamma-Gamma1.2891.27597.9%
2011-03-01Boosting1.7141.72996.1%
2011-06-01BG/NBD + Gamma-Gamma1.2021.09698.1%
2011-06-01Boosting1.3101.15092.0%
2011-09-01BG/NBD + Gamma-Gamma0.8050.67693.1%
2011-09-01Boosting1.0310.99687.8%

September favors boosting on count deviance: 1.043 versus BG/NBD 1.071. That reversal remains in the pooled comparison. Earlier final outcomes become eligible for later fits only when fully observed; no current-window outcomes enter training.

Nominal 90% BG/NBD count bands cover 97.8% of final observations, average width 2.41 purchase days; spend bands cover 96.2%, average width £1,237.62. Boosted count bands cover 98.4%, width 2.82; spend bands cover 91.8%, width £984.34. Wide, discrete intervals can overcover. Boosted spend coverage falls to 87.8% in September, so pooled coverage is not a guarantee.

Repeated-line accounting, returns and spend dependence

The multiset-union sensitivity keeps repeated within-sheet lines while removing duplicate sheet copies. All models refit with the frozen complexity. The primary difference remains 0.094 (95% interval 0.080–0.109), favoring BG/NBD.

Final signed net spend is £5,550,737.62. The same gross forecasts divided by this different outcome yield 0.986 for BG/NBD and 1.276 for boosting. This is a target-mismatch diagnostic, not a model that predicts returns or negative revenue.

The largest 1% of final customer-window spending accounts for 37.2% of gross value. Repeat-frequency/spend Spearman correlations range from 0.227 to 0.277, challenging Gamma-Gamma’s strict independence assumption. Large purchases remain in the analysis.

1,000 paired customer-cluster bootstrap draws with all final windows retained per sampled customer; 95% percentile intervals condition on fitted models. BG/NBD predictive distributions use 2,000 draws/customer; 90% challenger residual bands use separate December calibration customers.

04 / DATA

Resolve overlapping sheets before counting customers.

Online Retail II contains 1,067,371 lines from a UK non-store giftware retailer, with many wholesale customers. The two workbook sheets overlap in December 2010. Treating every line as distinct would double-count some purchases.

Exact complete-row deduplication leaves 1,033,036 lines; 235,151 lack customer ID and cannot enter longitudinal forecasts. Eligible positive-price, positive-quantity, noncancelled purchases produce 779,425 lines, 36,969 invoices, 33,107 customer purchase days and 5,878 customers. Combining same-day invoices prevents invoice splits from becoming artificial repeat purchases.

Primary gross positive spend is £17,374,804.27 in nominal pounds. Credits and signed quantities are retained separately for the net-spend sensitivity. Credit invoices do not reliably identify the original purchase, and the first observed purchase is not a verified acquisition date. Customer IDs join histories but never enter predictors; raw records and individual predictions are not published.

Source: Chen, D. (2012), Online Retail II, UCI Machine Learning Repository, doi:10.24432/C5CG6D. Licensed CC BY 4.0. Original aggregate analysis is published with attribution; no publisher endorsement is implied.

05 / METHOD & LIMITATIONS

Freeze selection and respect when outcomes become known.

Development snapshots begin March, June and September 2010. December validation customers split by stable hash: half select the boosted model’s complexity, half calibrate predictive bands. Seven leaves beat fifteen on validation deviance (1.501 versus 1.532). Final rolling fits use only completed labels and never retune on the current forecast window.

BG/NBD estimates heterogeneous purchase rates and latent dropout from prior histories. Gamma-Gamma pools repeat purchase-day spend; customers without repeats use its population mean. All likelihood starts converge without parameter-bound hits. The challenger combines probability of any purchase with one plus expected excess count and a conditional spend model. The RFM and recent-purchase rules remain visible baselines.

Model time starts at first observation, not proven acquisition. BG/NBD expected counts use an integral valid when its dropout-shape parameter is below one, checked against higher-order quadrature and simulation. Predictive distributions condition on estimated population parameters. The challenger’s empirical bands use reserved validation customers; seasonal drift and repeated customers prevent a guaranteed coverage claim.

  • Older single-retailer giftware and wholesale behavior may not transfer to a current consumer business.
  • Missing IDs, overlapping sheets, repeated line ambiguity and cancellations materially affect accounting.
  • Positive-purchase spend excludes returns and is not profit or net realized revenue.
  • Seasonality, refitting and customer dependence undermine guaranteed predictive interval coverage.
  • Only three primary final windows are observed; unseen future shocks and parameter uncertainty remain unresolved.
  • No randomized win-back intervention or observed permanent churn label is available.
  • Independent technical review is pending.

Method references: BG/NBD derivation, expectation correction and Gamma-Gamma monetary model.

Frozen customer-base protocol ↗

06 / CODE

Reproduce the purchasing and forecast accounting.

Run S03-2f7d21cd-572e3627
Analysis commit 2f7d21cdbdd1255931a9297ca19303f3551e421e

git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S03/study.py
uv run python scripts/report_s03.py
uv run python -W error -m unittest discover -s tests -v

Checks cover overlapping-sheet accounting, positive purchases versus returns, feature and label timing, probability identities, numerical integration, predictive sampling, metrics and display accounting. Independent technical review is pending.