← Research catalog

S31 / INSURANCE

Full-loss evidence leaves the model upgrade unresolved.

Boosting’s small full-claim accuracy improvement over interpretable frequency–severity regression is inconclusive. Tail losses, segment calibration and geographic transfer change the picture.

Evaluated public-data studyResearch code ↗Executive summary ↓

136,271 final policies · 5,426 claims · 72,158.9 policy-years · historical French motor insurance, approximately 2011–2013

01 / DECISION · EXECUTIVE SUMMARY

Require tail and calibration evidence before replacing the pricing model.

An insurance leader needs a model that estimates expected claims reliably across the portfolio. On this historical policy holdout, the more flexible model’s slight accuracy gain does not establish a dependable replacement case. The largest losses remain a material source of uncertainty.

€147.76observed loss per policy-year
€138.35boosted estimate per policy-year
38.0%source loss in the largest 1% of claims

Keep the interpretable model in the comparison.

Full-loss prediction error is 77.894 for boosting versus 78.424 for Poisson–Gamma regression. The paired difference is -0.531, with a 95% interval of -1.626 to +0.811. This does not establish a clear advantage. The error measure is Tweedie deviance; lower is better, and it is not a percentage saving.

A loading cannot fix the loss estimate.

Boosting predicts about 6.4% less total loss than observed; the interpretable model predicts about 7.4% more. Both calibration intervals include an exact match. Expense assumptions change an illustrative premium, but they do not make the underlying claim estimate more reliable.

02 / IMPLICATION

Make the loss definition part of the decision.

When each claim is capped at €50,000, the boosted model’s advantage is clearer. That narrower comparison answers a different question. Large claims are part of the full-loss target and cannot be removed merely because they make the result less favorable.

A reserved-region test also favors boosting relative to regression, yet boosting still overpredicts that region’s recorded loss by 19.1%. Relative improvement and absolute calibration are separate requirements. The evidence supports continued model comparison and prospective validation; no current quote, rate filing or realized savings is claimed.

03 / EVIDENCE

Compare models on the same untouched policies.

The final holdout contains 136,271 policies, 5,426 claims and 5,114 policies with at least one claim. All claims from a policy remain together. Primary deviance uses exposure weights and fixed power 1.5.

Full recorded loss · held-out policies · lower deviance is better
ModelTweedie deviance95% intervalPredicted / observed loss
Training portfolio mean81.23377.122–86.2771.164
Poisson–Gamma regression78.42473.706–84.4631.074
Direct Tweedie regression78.75973.859–85.1421.135
Boosted frequency–severity77.89472.738–84.7170.936

At portfolio level, boosting estimates 7.33 claims per 100 policy-years and €1,886.53 per predicted claim, versus 7.52 observed claims and €1,965.01 observed average claim size. The frequency–severity decomposition reconciles to total predicted loss.

HISTORICAL LOSS EXPLORER

Compare the expected loss with the loss actually recorded.

Model, loss basis and cohort controls select saved final results. Expense loading is a hypothetical calculation; it does not change the measured model comparison.

136,271 final policies · 5,426 claims · 72,158.9 policy-years · full recorded loss

€138.35estimated loss per policy-year
€147.76observed loss per policy-year
0.94×predicted / observed total loss

Calibration ratio 95% interval: 0.83–1.06. A ratio of 1 matches recorded total loss; below 1 underpredicts it. These intervals describe the observed sample, conditional on the fitted model.

Separate claim frequency from claim size.

Observed: 7.52 claims per 100 policy-years × €1,965.01 per claim.

Estimated: 7.33 claims per 100 policy-years × €1,886.53 per claim. Severity is weighted by predicted claim frequency across this cohort.

Illustrative expense calculation

€138.35 ÷ (1 − 0.20) = €172.93 per policy-year.

The 20% expense share is assumed. This is not an insurance quote or proposed rate; profit, taxes, capital, reinsurance and regulatory constraints are excluded. A loading cannot repair a miscalibrated loss estimate.

Predicted to observed loss ratios and 95 percent policy bootstrap intervals for the selected cohortPortfolio meanPoisson–GammaDirect TweedieBoosted model0.0×1.0×2.1×
All policies: dots show predicted/observed total loss; lines show paired-policy bootstrap intervals. The dashed line at 1 is exact portfolio calibration. Wide intervals are part of the evidence.
All policies · full recorded claims
ModelEstimated €/policy-yearPredicted / observed95% interval
Training portfolio mean€172.001.1641.029–1.321
Poisson–Gamma regression€158.741.0740.954–1.220
Direct Tweedie regression€167.631.1351.010–1.286
Boosted frequency–severity€138.350.9360.833–1.063
Portfolio risk-group calibration for the selected model

Ten equal-count score groups use the entire final portfolio, independent of the cohort selector above. They describe calibration; they are not ten independent experiments.

Boosted frequency–severity · uncapped loss · whole final portfolio
Risk groupPoliciesClaimsEstimated €/yearObserved €/yearRatio
113,628320€55.99€120.920.46
213,627387€72.74€83.570.87
313,627376€83.63€79.081.06
413,627422€93.83€87.761.07
513,627405€105.35€101.621.04
613,627441€118.98€106.201.12
713,627458€137.05€128.361.07
813,627620€164.50€196.500.84
913,627687€215.16€237.070.91
1013,6271,310€444.48€433.001.03

Cohorts require at least 500 final policies. Empirical intervals omit unseen catastrophic losses and unobserved dependence. The capped analysis changes the target; it does not establish the full-loss model’s performance.

Per-claim cap sensitivity and geographic transfer

€50,000 per-claim sensitivity

The capped-target deviance difference is -0.721 (95% interval -1.111 to -0.317), favoring boosting. All models are refitted against capped outcomes using their uncapped development-selected settings. Capped observed loss is €130.75 per policy-year; boosting predicts €123.93. This is a separate estimand, not a contractual coverage limit.

Reserved Ile-de-France policies

All 69,789 policies and 2,591 claims in this region are withheld from geographic-stress fitting and selection. Both models omit Region as a predictor and fit 485,915 non-reserved policies. The paired deviance difference is -4.567 (95% interval -6.017 to -2.791).

Separate region and independently selected models · full claim amounts
ModelDeviancePredicted / observed loss95% ratio interval
Poisson–Gamma regression83.0101.8421.615–2.123
Boosted frequency–severity78.4431.1911.045–1.374

The better relative score still accompanies overprediction. A new-region test cannot establish future-time performance.

Tail concentration, segment gaps and count dispersion

Across the source, 88 claims above €50,000 account for €17.92 million of loss; the single largest claim is €4.08 million, or 6.8% of all loss. Capping changes the evidence materially.

The boosted 60+ driver cohort has a predicted/observed ratio of 0.721 (95% interval 0.546–1.005), on 22,637 policies and 887 claims. Area F has ratio 0.604 (0.315–1.311), on 3,628 policies and 150 claims. These descriptive slices have wide uncertainty and do not establish causal effects or fairness.

Mean squared Pearson count residuals are 1.778 for Poisson regression and 1.741 for boosting, above the Poisson reference of 1. The analysis uses empirical policy bootstrap intervals; it does not claim Poisson claim-count prediction intervals.

Paired policy bootstrap: 500 final and geographic-stress draws, 250 per segment; 95% percentile intervals conditional on fitted models and the observed empirical loss tail. No independent-year or catastrophe uncertainty claim. The metrics use 500 draws and the explorer’s segment tables use 250; their portfolio-ratio intervals can differ slightly from Monte Carlo variation.

04 / DATA

Reconcile every claim to its exposure.

The archived CASdatasets 1.2-0 files contain 677,991 policies and 26,444 claims from an unidentified French motor insurer. Every claim joins to a policy and every policy’s recorded claim count agrees with its severity rows. There are no missing source fields.

The analysis retains 235 identical severity rows because equal settlement amounts can represent distinct claims; deleting them breaks count reconciliation. It also retains 1,224 exposures longer than one year, up to 2.01 years. Total exposure is 358,482.8 policy-years and total nominal claim loss is €59.91 million.

Policy IDs join the files and allocate splits. They never enter the model. Exposure is a denominator and weight; outcomes never enter the predictors. R category labels are preserved so internal factor codes cannot corrupt policy joins. Raw policies and individual predictions are not published.

The approximate source period is 2011–2013. There are no reliable policy dates, independently verified feature timestamps or claim-development histories. Some claim amounts reflect French IRSA/IDA settlement conventions. No temporal validation, ultimate-loss estimate or inflation adjustment is claimed.

Source: Dutang, Charpentier & Gallic, Insurance dataset, Recherche Data Gouv, version 1.1, July 12, 2024; CASdatasets 1.2-0 archive. Publisher dictionary. Dataset metadata uses Etalab Open Licence 2.0; package code is separately licensed GPL ≥2. Original aggregate analysis is presented with attribution; no publisher endorsement is implied.

05 / METHOD & LIMITATIONS

Separate model selection, full-loss evaluation and stress tests.

A fixed salted policy-ID hash allocates 405,817 development policies, 135,903 validation policies and 136,271 final policies. Model settings are selected on validation deviance, then refitted on 541,720 development/validation policies. The final cohort is never used for tuning or recalibration.

Poisson frequency uses exposure weights, equivalent to a log-exposure offset. Gamma severity uses claim-count weights. The interpretable models use training-fitted age splines, transformed density and bonus-malus, and risk categories. Direct Tweedie estimates the full pure-premium mean. Histogram boosting uses Poisson frequency and Gamma severity with the same exposure and claim-count identities.

Validation selects regularization .0001 for both regression families and seven leaves for boosting. All candidate comparisons remain in the result artifact. Bootstrap intervals condition on the fitted models and observed policies; shared drivers, catastrophic shocks and future market change remain unresolved.

  • Historical French motor data do not validate a current insurance price or rate filing.
  • No policy dates support a temporal holdout or ultimate-claim development model.
  • Nominal euros are not inflation-adjusted and some amounts reflect settlement conventions.
  • Heavy losses and unobserved dependence can make empirical bootstrap intervals too narrow for future risk.
  • Driver age and geography are descriptive benchmark features, not causal effects or a fairness certification.
  • No policyholder eligibility or pricing action is taken; expense scenarios exclude taxes, profit, capital and regulatory constraints.
  • Independent technical review is pending.
Frozen actuarial, tail and geographic protocol ↗

06 / CODE

Reproduce the loss and exposure accounting.

Run S31-02c309de-1e141502
Analysis commit 02c309de271f6bb119a598b0c5972a3a697e553b

git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S31/study.py
uv run python scripts/report_s31.py
uv run python -W error -m unittest discover -s tests -v

Pinned publisher archive and extracted-file hashes, locked dependencies and fixed seeds. Checks cover policy joins, factor labels, exposure and severity identities, split stability, tail definitions, uncertainty and display accounting. Independent technical review is pending.