01 / DECISION · EXECUTIVE SUMMARY
Keep the simple benchmark until the upgrade proves itself.
A sales leader with limited calling capacity needs a reliable way to prioritize contacts. In this historical bank dataset, the rule based on previous campaign success and contact history performs better than the logistic model selected on earlier records. Greater model complexity does not earn its place in this test.
What a leader should do next
Use the rule as a benchmark in a current, controlled validation. Require a replacement to improve both forecast quality and the decision at the actual contact capacity. Monitor changing response rates before scaling.
What this does not establish
Every record comes from an observed campaign contact. The study predicts recorded response; it cannot tell how many subscriptions the call caused. There is no measured profit or current-client outcome. The historical record sequence also lacks exact dates and customer identifiers.
02 / IMPLICATION
A better ranking score is not the whole decision.
The selected logistic model has a higher AUROC than the rule, yet identifies fewer subscriptions at the specified capacity and has worse probability forecasts. Compare the metric that serves the decision, and keep adverse findings visible.
The primary log-loss difference (selected model minus rule) is 0.0822 natural-log units, with a 95% interval of 0.0470–0.1177. Positive values favor the simpler rule. This advantage persists across the prespecified bootstrap block-size checks.
03 / EVIDENCE
Measured comparisons, explicit assumptions.
Final holdout: 8,238 source-ordered contacts and 2,540 subscriptions. Primary metric: log loss, which penalizes confident wrong probabilities; lower is better.
| Model | Log loss | 95% interval | AUROC |
|---|---|---|---|
| Random allocation | 0.7423 | 0.6631–0.8209 | 0.500 |
| History and recency rule | 0.6349 | 0.5730–0.6938 | 0.672 |
| Regularized logistic · selected in development | 0.7171 | 0.6392–0.7909 | 0.692 |
| Additive spline logistic | 0.7536 | 0.6704–0.8332 | 0.573 |
| Gradient boosting | 0.7448 | 0.6635–0.8228 | 0.633 |
| Boosting with richer attributes | 0.7284 | 0.6523–0.8023 | 0.680 |
Richer-attribute boosting is a sensitivity analysis. Random allocation uses the preceding calibration cohort’s response rate for its constant probability forecast. All models were fixed before final scoring.
CONTACT CAPACITY EXPLORER
Which recorded responses would the ranking capture?
Choose a saved model and capacity. The chart and counts use the untouched historical holdout; controls do not retrain models or change their measured accuracy.
Conditional 95% precision interval: 48.7%–65.1%. Captures 36.8% of recorded subscriptions; interval 32.4%–40.9%. Intervals describe resampling uncertainty, not uncertainty about already observed counts.
| Model | Responses | Precision | Captured |
|---|---|---|---|
| Random allocation | 507.8 expected | 30.8% | 20.0% |
| History and recency rule | 934 | 56.7% | 36.8% |
| Regularized logistic | 843 | 51.2% | 33.2% |
| Additive spline logistic | 790 | 48.0% | 31.1% |
| Gradient boosting | 876 | 53.2% | 34.5% |
| Boosting with richer attributes | 831 | 50.5% | 32.7% |
ASSUMED ECONOMICS / SCENARIO ONLY
What would these counts imply under your assumptions?
Hypothetical USD-equivalent units; neither cost nor response value is observed in the data. No incremental profit is estimated.
$76,930 illustrative value less contact cost
Conditional scenario range: $63,736–$90,791. Assumed break-even response value: $18.
Calculation: selected responses × assumed response value − selected contacts × assumed contact cost. Range rescales the conditional precision interval at the selected count. It excludes retraining, future population shifts and uncertainty in your economic assumptions.
Chronological failure groups and probability calibration
These are four consecutive groups of final source rows, not calendar quarters. Response rates shift substantially; earlier calibration does not protect against later change.
| Group | Contacts | Response rate | Rule loss | Model loss |
|---|---|---|---|---|
| 1 | 2060 | 8.0% | 0.2803 | 0.2725 |
| 2 | 2060 | 23.5% | 0.5795 | 0.6323 |
| 3 | 2059 | 39.8% | 0.7968 | 0.8802 |
| 4 | 2059 | 52.1% | 0.8831 | 1.0836 |
| Contacts | Mean forecast | Observed rate |
|---|---|---|
| 5276 | 7.5% | 20.6% |
| 1830 | 12.9% | 48.4% |
| 344 | 24.8% | 51.7% |
| 529 | 36.3% | 44.8% |
| 221 | 43.2% | 58.8% |
| 24 | 54.1% | 62.5% |
| 10 | 63.0% | 50.0% |
| 3 | 71.0% | 33.3% |
| 1 | 87.7% | 100.0% |
Adding demographic and loan/default attributes does not reverse the primary conclusion. Removing two test profiles seen in earlier data also leaves the comparison intact. Matching anonymized profiles do not prove matching people. Full diagnostics and sensitivity report ↗
500 paired circular moving-block bootstrap draws, 100 source-ordered test records per block; conditional on fixed fitted models. Block lengths 50/200 are sensitivity checks. These intervals do not remove dependence from unidentified repeated customers or include model-retraining uncertainty.
04 / DATA
Historical contacts, with a strict pre-call cutoff.
Moro, Rita & Cortez (2014), Bank Marketing, UCI, doi:10.24432/C5K306. The full bank-additional file contains 41,188 rows, ordered by date by its publisher. Coverage is May 2008–November 2010; exact per-call dates and stable customer IDs are unavailable.
Development uses the first 28,831 rows; probability calibration uses the next 4,119; evaluation uses the final 8,238. There are 12 exact repeated rows, two final profiles seen in earlier records, and 570 final contacts with a month category not seen in development. Missing categories remain explicit.
Call duration is excluded because it is unavailable before the call. All five economic indicators are excluded because their historical release vintages cannot be verified. Current-campaign call count is reduced by one to exclude the current call.
Publisher, data and CC BY 4.0 terms ↗
No individual contact records are published here. Aggregate exports retain data/source hashes and the observed sampling boundaries.
05 / METHOD & LIMITATIONS
Chronological evaluation before model preference.
Two expanding development folds choose regularization or tree complexity using log loss. Logistic regression, additive splines and histogram gradient boosting train on the first 70% of records. A disjoint next 10% fits sigmoid probability calibration. The last 20% is untouched until evaluation.
The principal family is selected only from development results. All candidates remain in the final comparison. The simple rule prioritizes previous success, prior contact, recency and fewer earlier calls. The primary model uses operational contact history; a separate richer-feature sensitivity adds demographic and loan attributes.
- Single bank, historical observed-contact population; no randomized no-contact comparison.
- No stable customer IDs or exact call dates; dependence and repeated contacts cannot be fully resolved.
- Calibrated probabilities can fail under later campaign shifts. Bootstrap intervals do not include model-retraining uncertainty.
- Economic indicator release vintages cannot be verified, so all five are excluded.
- Matching anonymized profiles do not establish matching people. Independent technical review is pending.
At 20% capacity the rule’s precision is 56.7%, with a conditional 95% interval of 48.7%–65.1%. Decision intervals resample the frozen selected indicators; resampled selection counts may vary. They are not causal intervals or guarantees for a new campaign.
06 / CODE
Follow the finding back to the run.
Run S28-da76dd82-74adfc57
Analysis commit da76dd823b1b242cd90fc0280f93926d2bd575ba
git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error -m unittest discover -s tests -v
uv run python -W error studies/S28/study.py
uv run python scripts/report_s28.pyPython 3.11.16 · locked dependencies · fixed seeds · 15 correctness checks. Source and CSV checksums are included in the result export. Independent technical review is pending.