← Research catalog

S47 / TELECOMMUNICATIONS

Service friction identifies risk. Can recovery change it?

Boosting improves three-month churn forecasts and finds 90 of 94 recorded churn outcomes within a 124-record queue. A randomized service test is still needed to measure retention benefit.

Evaluated public-data studyResearch code ↗Executive summary ↓

Single-company Iranian telecom sample · 3,150 source rows · 624 final records / 561 profiles · publisher-described three-month planning gap

01 / DECISION & EXECUTIVE SUMMARY

Use service signals to prioritize a test of recovery.

A service leader wants advance warning of departure and a defensible way to allocate limited attention. In this small historical sample, boosting forecasts churn more accurately than additive regression, even after excluding Status and calculated customer value. Complaints and call failures contribute useful predictive information.

90 / 94observed churn outcomes identified at 20% capacity
124 recordsselected from 624 · 34 selected records did not churn
0.0981 vs 0.1576boosting vs additive log loss · lower is better

Predictive evidence supports further testing.

The paired log-loss difference is -0.0595 (95% interval -0.0826 to -0.0347), favoring boosting. Its queue misses four recorded churn outcomes; additive regression misses nine, and the complaint/low-usage rule misses 37 at the same capacity.

No churn outcome is claimed prevented.

The file has no customer ID or dates and includes repeated profiles with conflicting labels. Identical profiles stay together in evaluation, but the split is not verified at the individual-customer level and does not test a later time period. Risk ranking cannot establish that resolving a complaint causes retention.

02 / IMPLICATION

Keep the service signals; measure the intervention separately.

Removing complaints and call failures raises boosted log loss by 0.0511 (95% interval 0.0255–0.0774). Adding the uncertain Status or calculated-value fields yields no clear further gain. Before deploying a queue, verify field timing and customer identity on a later cohort, then randomize the recovery offer.

For planning, an assumed reduction from 15% to 12% churn requires approximately 2,036 participants per arm—4,072 total—under the stated simple trial assumptions. That exceeds this benchmark’s entire sample. The historical model supplies neither a treatment-effect estimate nor a validated campaign budget.

03 / EVIDENCE

Separate ranking quality, probability calibration and treatment value.

The fixed final partition has 624 records, 561 distinct primary-feature profiles and 94 churn outcomes. The primary comparison uses calibrated log loss and 1,000 paired profile-bootstrap draws. Identical primary-feature profiles stay together in every split and development fold.

Core features · Status, calculated value and redundant Age excluded
ModelLog loss95% intervalAverage precisionChurn found / 124 contacts
Development prevalence0.42410.3723–0.47790.150619 / 94
Complaint / low-usage rule0.30110.2493–0.35020.506357 / 94
Ridge logistic0.20230.1624–0.24240.784474 / 94
Additive logistic0.15760.1245–0.19360.862285 / 94
Histogram boosting0.09810.0703–0.13330.949090 / 94

At 20% capacity, boosting’s recall is 95.7% (95% interval 89.9–100.0%) and precision is 72.6%. Average precision is 0.949 (0.906–0.974) versus additive 0.862 (0.792–0.923); the observed churn prevalence is 0.151. These are historical classification results.

HISTORICAL SERVICE QUEUE

See who was at risk before designing the recovery test.

Model and feature controls select saved evaluations. Capacity changes the number of historically ranked records; no churn outcome is claimed prevented.

90 / 94observed churn outcomes found
4churn outcomes outside the queue
34selected records that did not churn

Recall 95.7% · 95% profile-bootstrap interval 89.9%–100.0%. Precision 72.6%. These describe recorded outcomes under historical ranking.

Model log loss: 0.0981 (95% interval 0.0703–0.1333; lower is better). Average precision: 0.9490 (0.9064–0.9742; higher is better). These metrics remain fixed when capacity changes.

Observed churn recall across saved service capacities for the selected model and feature set0%0%50%50%100%100%Service capacity · percent of final records
Vertical axis: fraction of observed churn identified. Thin lines: 95% profile-bootstrap intervals. Copper dot: selected capacity. No intervention is simulated by this curve.
20% capacity · Core features · status/value excluded
ModelChurn foundMissedPrecision
Additive85968.5%
Boosting90472.6%
Ridge logistic742059.7%
Prevalence197515.3%
Complaint / usage rule573746.0%
Probability calibration for the selected model and features

Ten equal-count score groups compare predicted churn with the rate actually observed. Intervals resample profiles within each bin; sparse-bin intervals are descriptive and may be unstable. A zero-event bin produces a zero bootstrap interval, which cannot rule out future churn; percentages are rounded to one decimal place.

Whole final partition · independent of selected capacity
Risk groupRowsChurnPredictedObserved95% observed interval
16300.0%0.0%0.0%–0.0%
26300.0%0.0%0.0%–0.0%
36300.1%0.0%0.0%–0.0%
46300.1%0.0%0.0%–0.0%
56200.2%0.0%0.0%–0.0%
66200.6%0.0%0.0%–0.0%
76211.7%1.6%0.0%–4.8%
86238.7%4.8%0.0%–10.7%
9623159.6%50.0%38.3%–62.3%
10625996.4%95.2%89.2%–100.0%

Inspect the additive model’s associations.

This separate view always uses the core additive fit and its 100 development-bootstrap refits. The selector changes the feature contrast; the queue controls above do not alter it.

Complains: 0 → 1, with all other fields held at their recorded development values.

Mean model-risk difference: +36.7 percentage points; 95% stability interval 28.5 to 45.3. The contrast is positive in 100.0% of 100 refits.

Selected additive model risk contrast and bootstrap stability interval-50 pp-25 pp0 pp25 pp50 pp
The contrast describes a fitted model. Correlated features and unusual combinations can change its interpretation; it does not estimate the effect of resolving a complaint or changing call behavior.

Plan the experiment needed to measure retention benefit.

Proposed test: pre-register eligibility, randomize usual service versus an additional recovery offer 1:1 within pre-contact risk strata, and measure intention-to-treat churn at three months. Track contact costs and adverse outcomes. Persistent identities, complete follow-up and protection against cross-arm contamination are prerequisites. No trial has been run.

Prospective sample-size scenario

15% → 12% assumed churn requires approximately 2,036 participants per arm, or 4,072 total.

Two-sided alpha 5%, power 80%, equal independent groups. Normal approximation; attrition, clustering, noncompliance and multiple tests are excluded. The assumed reduction is not inferred from the historical risk model.

Feature sensitivities and calibration tradeoffs
Boosted variant minus core boosted log loss · negative favors the variant
ChangeDifference95% paired interval
Remove service friction0.05110.0255 to 0.0774
Add Status-0.0082-0.0218 to 0.0044
Add calculated value0.0018-0.0053 to 0.0093
Add Status and value-0.0042-0.0177 to 0.0086

All variants use the same partitions and select their settings inside development. The added derived fields are sensitivities, not a new primary model selected after seeing final outcomes.

Reserved probability calibration lowers boosted log loss from 0.1002 to 0.0981, but increases its predicted/observed churn ratio from 1.017 to 1.104 (95% calibrated interval 1.016–1.204). Additive log loss slightly worsens from 0.1575 to 0.1576. Better ranking or log loss does not guarantee accurate total-event prediction.

Giving each repeated profile total weight one preserves the ordering: boosted log loss 0.0868 versus additive 0.1425. This is a different weighting estimand, not a solution to missing identities. Age groups 1 and 5 and the contractual-tariff group fail the minimum 30-record/five-events-per-class rule for public subgroup summaries; they remain in totals.

Why the association is not an intervention effect

The core additive complaint contrast is +36.7 percentage points (95% stability interval 28.5–45.3); the call-failure contrast, from one to twelve failures, is +12.2 points (6.8–17.3). Both signs are positive in all 100 development refits. These contrasts hold correlated features fixed and can represent unusual combinations.

Hypothesized common causes of service signals and churn, distinct from the proposed randomized recovery effectLatent service qualityLatent engagementFailures / complaintsObserved usageLater churnArrows are hypotheses; randomized recovery is the proposed test.
Shared causes can produce both service signals and departure. The full report supplies the causal diagram and prospective trial plan. No arrow is identified from this observational benchmark.

Holding call count fixed while increasing seconds yields a positive model-risk contrast; increasing call count at fixed seconds yields a negative one. These conditional patterns are not advice to increase calls or reduce duration.

1,000 paired primary-profile bootstrap draws; 95% percentile intervals condition on the fitted model. Capacity ranking is recomputed per draw. Additive effect stability refits 100 profile-bootstrap development samples using a fixed calibration map.

04 / DATA

Keep repeated profiles together and preserve the planning gap.

The official Iranian Churn file has 3,150 records, 14 columns, 495 churn labels and no missing cells. The publisher describes first-nine-month feature aggregates and churn status at month twelve. No actual date or anonymous customer ID appears in the file, despite a dictionary reference to ID.

The ten-field primary representation creates 2,826 distinct profiles and 324 repeated-profile rows. Twenty-two groups contain conflicting churn labels, totaling 83 rows. All records are retained, grouped across partitions and resampled as whole profiles. No assertion of verified independent people is made.

Age contains only five representative values mapped exactly from age group. It is excluded as redundant, while age group remains categorical. Seven charge values exceed the dictionary’s stated maximum and are retained with the discrepancy documented. Complains is a binary flag; no complaint count, recent-activity date or usage trend can be reconstructed.

Status and calculated Customer Value are excluded from primary inputs. Their operational definitions are too limited to establish safe decision-time availability; they are not proven leakage. Customer Value has no verified currency/profit interpretation. Raw records, profile hashes and individual predictions are not published.

Source: Iranian Churn, UCI Machine Learning Repository (2020), doi:10.24432/C5JW3Z, licensed CC BY 4.0. Original aggregate analysis is presented with attribution; no publisher endorsement is implied.

05 / METHOD & LIMITATIONS

Select within development, calibrate separately, then evaluate.

A fixed five-fold grouped stratification assigns 1,897 development records, 629 calibration records and 624 final records. Core additive and boosted models receive three-fold outer development checks with three inner selection folds. All final feature variants select their settings within development only.

Additive regression uses training-fitted cubic splines of log-transformed numeric inputs and one-hot categories. Ridge logistic omits the splines. Histogram boosting uses raw numeric and categorical inputs. All variants select C=10 for additive regression and fifteen leaves for boosting; core linear C=10. Probabilities are calibrated using only the reserved partition. No class weights or final-label threshold optimization are applied.

Association stability refits 100 development-profile samples at fixed chosen regularization and uses the original calibration map. It does not quantify a causal effect. The planning calculator uses an equal-group normal approximation with two-sided alpha .05 and power .80; saved values agree with the independent two-proportion sample-size calculation.

  • Small single-company sample with unspecified collection year and no temporal replication.
  • No actual customer ID; repetitions and conflicting identical profiles prevent verified customer independence.
  • Status and calculated customer value have insufficiently documented semantics; their sensitivity gains do not prove operational availability.
  • Age is a category representative, and charge amount includes seven values beyond the dictionary range.
  • Associations and model-risk contrasts do not identify benefits from resolving complaints or changing usage.
  • Sparse contractual churn and correlated features limit subgroup and effect interpretation.
  • Independent technical review is pending.
Frozen service-friction protocol ↗

06 / CODE

Reproduce the comparison and its limits.

Run S47-a0ada992-696c3a18
Analysis commit a0ada99288b53cae3fd494991e16e055c7c5c39a

git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S47/study.py
uv run python scripts/report_s47.py
uv run python -W error -m unittest discover -s tests -v

Checks cover grouped splits, excluded predictors, fold-local transformations, score arithmetic, ties and capacity accounting, cluster resampling, sample-size feasibility and display identities. Independent technical review is pending.