01 / DECISION & EXECUTIVE SUMMARY
Use service signals to prioritize a test of recovery.
A service leader wants advance warning of departure and a defensible way to allocate limited attention. In this small historical sample, boosting forecasts churn more accurately than additive regression, even after excluding Status and calculated customer value. Complaints and call failures contribute useful predictive information.
Predictive evidence supports further testing.
The paired log-loss difference is -0.0595 (95% interval -0.0826 to -0.0347), favoring boosting. Its queue misses four recorded churn outcomes; additive regression misses nine, and the complaint/low-usage rule misses 37 at the same capacity.
No churn outcome is claimed prevented.
The file has no customer ID or dates and includes repeated profiles with conflicting labels. Identical profiles stay together in evaluation, but the split is not verified at the individual-customer level and does not test a later time period. Risk ranking cannot establish that resolving a complaint causes retention.
02 / IMPLICATION
Keep the service signals; measure the intervention separately.
Removing complaints and call failures raises boosted log loss by 0.0511 (95% interval 0.0255–0.0774). Adding the uncertain Status or calculated-value fields yields no clear further gain. Before deploying a queue, verify field timing and customer identity on a later cohort, then randomize the recovery offer.
For planning, an assumed reduction from 15% to 12% churn requires approximately 2,036 participants per arm—4,072 total—under the stated simple trial assumptions. That exceeds this benchmark’s entire sample. The historical model supplies neither a treatment-effect estimate nor a validated campaign budget.
03 / EVIDENCE
Separate ranking quality, probability calibration and treatment value.
The fixed final partition has 624 records, 561 distinct primary-feature profiles and 94 churn outcomes. The primary comparison uses calibrated log loss and 1,000 paired profile-bootstrap draws. Identical primary-feature profiles stay together in every split and development fold.
| Model | Log loss | 95% interval | Average precision | Churn found / 124 contacts |
|---|---|---|---|---|
| Development prevalence | 0.4241 | 0.3723–0.4779 | 0.1506 | 19 / 94 |
| Complaint / low-usage rule | 0.3011 | 0.2493–0.3502 | 0.5063 | 57 / 94 |
| Ridge logistic | 0.2023 | 0.1624–0.2424 | 0.7844 | 74 / 94 |
| Additive logistic | 0.1576 | 0.1245–0.1936 | 0.8622 | 85 / 94 |
| Histogram boosting | 0.0981 | 0.0703–0.1333 | 0.9490 | 90 / 94 |
At 20% capacity, boosting’s recall is 95.7% (95% interval 89.9–100.0%) and precision is 72.6%. Average precision is 0.949 (0.906–0.974) versus additive 0.862 (0.792–0.923); the observed churn prevalence is 0.151. These are historical classification results.
HISTORICAL SERVICE QUEUE
See who was at risk before designing the recovery test.
Model and feature controls select saved evaluations. Capacity changes the number of historically ranked records; no churn outcome is claimed prevented.
Recall 95.7% · 95% profile-bootstrap interval 89.9%–100.0%. Precision 72.6%. These describe recorded outcomes under historical ranking.
Model log loss: 0.0981 (95% interval 0.0703–0.1333; lower is better). Average precision: 0.9490 (0.9064–0.9742; higher is better). These metrics remain fixed when capacity changes.
| Model | Churn found | Missed | Precision |
|---|---|---|---|
| Additive | 85 | 9 | 68.5% |
| Boosting | 90 | 4 | 72.6% |
| Ridge logistic | 74 | 20 | 59.7% |
| Prevalence | 19 | 75 | 15.3% |
| Complaint / usage rule | 57 | 37 | 46.0% |
Probability calibration for the selected model and features
Ten equal-count score groups compare predicted churn with the rate actually observed. Intervals resample profiles within each bin; sparse-bin intervals are descriptive and may be unstable. A zero-event bin produces a zero bootstrap interval, which cannot rule out future churn; percentages are rounded to one decimal place.
| Risk group | Rows | Churn | Predicted | Observed | 95% observed interval |
|---|---|---|---|---|---|
| 1 | 63 | 0 | 0.0% | 0.0% | 0.0%–0.0% |
| 2 | 63 | 0 | 0.0% | 0.0% | 0.0%–0.0% |
| 3 | 63 | 0 | 0.1% | 0.0% | 0.0%–0.0% |
| 4 | 63 | 0 | 0.1% | 0.0% | 0.0%–0.0% |
| 5 | 62 | 0 | 0.2% | 0.0% | 0.0%–0.0% |
| 6 | 62 | 0 | 0.6% | 0.0% | 0.0%–0.0% |
| 7 | 62 | 1 | 1.7% | 1.6% | 0.0%–4.8% |
| 8 | 62 | 3 | 8.7% | 4.8% | 0.0%–10.7% |
| 9 | 62 | 31 | 59.6% | 50.0% | 38.3%–62.3% |
| 10 | 62 | 59 | 96.4% | 95.2% | 89.2%–100.0% |
Inspect the additive model’s associations.
This separate view always uses the core additive fit and its 100 development-bootstrap refits. The selector changes the feature contrast; the queue controls above do not alter it.
Complains: 0 → 1, with all other fields held at their recorded development values.
Mean model-risk difference: +36.7 percentage points; 95% stability interval 28.5 to 45.3. The contrast is positive in 100.0% of 100 refits.
Plan the experiment needed to measure retention benefit.
Proposed test: pre-register eligibility, randomize usual service versus an additional recovery offer 1:1 within pre-contact risk strata, and measure intention-to-treat churn at three months. Track contact costs and adverse outcomes. Persistent identities, complete follow-up and protection against cross-arm contamination are prerequisites. No trial has been run.
Prospective sample-size scenario
15% → 12% assumed churn requires approximately 2,036 participants per arm, or 4,072 total.
Two-sided alpha 5%, power 80%, equal independent groups. Normal approximation; attrition, clustering, noncompliance and multiple tests are excluded. The assumed reduction is not inferred from the historical risk model.
Feature sensitivities and calibration tradeoffs
| Change | Difference | 95% paired interval |
|---|---|---|
| Remove service friction | 0.0511 | 0.0255 to 0.0774 |
| Add Status | -0.0082 | -0.0218 to 0.0044 |
| Add calculated value | 0.0018 | -0.0053 to 0.0093 |
| Add Status and value | -0.0042 | -0.0177 to 0.0086 |
All variants use the same partitions and select their settings inside development. The added derived fields are sensitivities, not a new primary model selected after seeing final outcomes.
Reserved probability calibration lowers boosted log loss from 0.1002 to 0.0981, but increases its predicted/observed churn ratio from 1.017 to 1.104 (95% calibrated interval 1.016–1.204). Additive log loss slightly worsens from 0.1575 to 0.1576. Better ranking or log loss does not guarantee accurate total-event prediction.
Giving each repeated profile total weight one preserves the ordering: boosted log loss 0.0868 versus additive 0.1425. This is a different weighting estimand, not a solution to missing identities. Age groups 1 and 5 and the contractual-tariff group fail the minimum 30-record/five-events-per-class rule for public subgroup summaries; they remain in totals.
Why the association is not an intervention effect
The core additive complaint contrast is +36.7 percentage points (95% stability interval 28.5–45.3); the call-failure contrast, from one to twelve failures, is +12.2 points (6.8–17.3). Both signs are positive in all 100 development refits. These contrasts hold correlated features fixed and can represent unusual combinations.
Holding call count fixed while increasing seconds yields a positive model-risk contrast; increasing call count at fixed seconds yields a negative one. These conditional patterns are not advice to increase calls or reduce duration.
1,000 paired primary-profile bootstrap draws; 95% percentile intervals condition on the fitted model. Capacity ranking is recomputed per draw. Additive effect stability refits 100 profile-bootstrap development samples using a fixed calibration map.
04 / DATA
Keep repeated profiles together and preserve the planning gap.
The official Iranian Churn file has 3,150 records, 14 columns, 495 churn labels and no missing cells. The publisher describes first-nine-month feature aggregates and churn status at month twelve. No actual date or anonymous customer ID appears in the file, despite a dictionary reference to ID.
The ten-field primary representation creates 2,826 distinct profiles and 324 repeated-profile rows. Twenty-two groups contain conflicting churn labels, totaling 83 rows. All records are retained, grouped across partitions and resampled as whole profiles. No assertion of verified independent people is made.
Age contains only five representative values mapped exactly from age group. It is excluded as redundant, while age group remains categorical. Seven charge values exceed the dictionary’s stated maximum and are retained with the discrepancy documented. Complains is a binary flag; no complaint count, recent-activity date or usage trend can be reconstructed.
Status and calculated Customer Value are excluded from primary inputs. Their operational definitions are too limited to establish safe decision-time availability; they are not proven leakage. Customer Value has no verified currency/profit interpretation. Raw records, profile hashes and individual predictions are not published.
Source: Iranian Churn, UCI Machine Learning Repository (2020), doi:10.24432/C5JW3Z, licensed CC BY 4.0. Original aggregate analysis is presented with attribution; no publisher endorsement is implied.
05 / METHOD & LIMITATIONS
Select within development, calibrate separately, then evaluate.
A fixed five-fold grouped stratification assigns 1,897 development records, 629 calibration records and 624 final records. Core additive and boosted models receive three-fold outer development checks with three inner selection folds. All final feature variants select their settings within development only.
Additive regression uses training-fitted cubic splines of log-transformed numeric inputs and one-hot categories. Ridge logistic omits the splines. Histogram boosting uses raw numeric and categorical inputs. All variants select C=10 for additive regression and fifteen leaves for boosting; core linear C=10. Probabilities are calibrated using only the reserved partition. No class weights or final-label threshold optimization are applied.
Association stability refits 100 development-profile samples at fixed chosen regularization and uses the original calibration map. It does not quantify a causal effect. The planning calculator uses an equal-group normal approximation with two-sided alpha .05 and power .80; saved values agree with the independent two-proportion sample-size calculation.
- Small single-company sample with unspecified collection year and no temporal replication.
- No actual customer ID; repetitions and conflicting identical profiles prevent verified customer independence.
- Status and calculated customer value have insufficiently documented semantics; their sensitivity gains do not prove operational availability.
- Age is a category representative, and charge amount includes seven values beyond the dictionary range.
- Associations and model-risk contrasts do not identify benefits from resolving complaints or changing usage.
- Sparse contractual churn and correlated features limit subgroup and effect interpretation.
- Independent technical review is pending.
06 / CODE
Reproduce the comparison and its limits.
Run S47-a0ada992-696c3a18
Analysis commit a0ada99288b53cae3fd494991e16e055c7c5c39a
git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S47/study.py
uv run python scripts/report_s47.py
uv run python -W error -m unittest discover -s tests -vChecks cover grouped splits, excluded predictors, fold-local transformations, score arithmetic, ties and capacity accounting, cluster resampling, sample-size feasibility and display identities. Independent technical review is pending.