01 / DECISION · EXECUTIVE SUMMARY
Require tail and calibration evidence before replacing the pricing model.
An insurance leader needs a model that estimates expected claims reliably across the portfolio. On this historical policy holdout, the more flexible model’s slight accuracy gain does not establish a dependable replacement case. The largest losses remain a material source of uncertainty.
Keep the interpretable model in the comparison.
Full-loss prediction error is 77.894 for boosting versus 78.424 for Poisson–Gamma regression. The paired difference is -0.531, with a 95% interval of -1.626 to +0.811. This does not establish a clear advantage. The error measure is Tweedie deviance; lower is better, and it is not a percentage saving.
A loading cannot fix the loss estimate.
Boosting predicts about 6.4% less total loss than observed; the interpretable model predicts about 7.4% more. Both calibration intervals include an exact match. Expense assumptions change an illustrative premium, but they do not make the underlying claim estimate more reliable.
02 / IMPLICATION
Make the loss definition part of the decision.
When each claim is capped at €50,000, the boosted model’s advantage is clearer. That narrower comparison answers a different question. Large claims are part of the full-loss target and cannot be removed merely because they make the result less favorable.
A reserved-region test also favors boosting relative to regression, yet boosting still overpredicts that region’s recorded loss by 19.1%. Relative improvement and absolute calibration are separate requirements. The evidence supports continued model comparison and prospective validation; no current quote, rate filing or realized savings is claimed.
03 / EVIDENCE
Compare models on the same untouched policies.
The final holdout contains 136,271 policies, 5,426 claims and 5,114 policies with at least one claim. All claims from a policy remain together. Primary deviance uses exposure weights and fixed power 1.5.
| Model | Tweedie deviance | 95% interval | Predicted / observed loss |
|---|---|---|---|
| Training portfolio mean | 81.233 | 77.122–86.277 | 1.164 |
| Poisson–Gamma regression | 78.424 | 73.706–84.463 | 1.074 |
| Direct Tweedie regression | 78.759 | 73.859–85.142 | 1.135 |
| Boosted frequency–severity | 77.894 | 72.738–84.717 | 0.936 |
At portfolio level, boosting estimates 7.33 claims per 100 policy-years and €1,886.53 per predicted claim, versus 7.52 observed claims and €1,965.01 observed average claim size. The frequency–severity decomposition reconciles to total predicted loss.
HISTORICAL LOSS EXPLORER
Compare the expected loss with the loss actually recorded.
Model, loss basis and cohort controls select saved final results. Expense loading is a hypothetical calculation; it does not change the measured model comparison.
136,271 final policies · 5,426 claims · 72,158.9 policy-years · full recorded loss
Calibration ratio 95% interval: 0.83–1.06. A ratio of 1 matches recorded total loss; below 1 underpredicts it. These intervals describe the observed sample, conditional on the fitted model.
Separate claim frequency from claim size.
Observed: 7.52 claims per 100 policy-years × €1,965.01 per claim.
Estimated: 7.33 claims per 100 policy-years × €1,886.53 per claim. Severity is weighted by predicted claim frequency across this cohort.
Illustrative expense calculation
€138.35 ÷ (1 − 0.20) = €172.93 per policy-year.
The 20% expense share is assumed. This is not an insurance quote or proposed rate; profit, taxes, capital, reinsurance and regulatory constraints are excluded. A loading cannot repair a miscalibrated loss estimate.
| Model | Estimated €/policy-year | Predicted / observed | 95% interval |
|---|---|---|---|
| Training portfolio mean | €172.00 | 1.164 | 1.029–1.321 |
| Poisson–Gamma regression | €158.74 | 1.074 | 0.954–1.220 |
| Direct Tweedie regression | €167.63 | 1.135 | 1.010–1.286 |
| Boosted frequency–severity | €138.35 | 0.936 | 0.833–1.063 |
Portfolio risk-group calibration for the selected model
Ten equal-count score groups use the entire final portfolio, independent of the cohort selector above. They describe calibration; they are not ten independent experiments.
| Risk group | Policies | Claims | Estimated €/year | Observed €/year | Ratio |
|---|---|---|---|---|---|
| 1 | 13,628 | 320 | €55.99 | €120.92 | 0.46 |
| 2 | 13,627 | 387 | €72.74 | €83.57 | 0.87 |
| 3 | 13,627 | 376 | €83.63 | €79.08 | 1.06 |
| 4 | 13,627 | 422 | €93.83 | €87.76 | 1.07 |
| 5 | 13,627 | 405 | €105.35 | €101.62 | 1.04 |
| 6 | 13,627 | 441 | €118.98 | €106.20 | 1.12 |
| 7 | 13,627 | 458 | €137.05 | €128.36 | 1.07 |
| 8 | 13,627 | 620 | €164.50 | €196.50 | 0.84 |
| 9 | 13,627 | 687 | €215.16 | €237.07 | 0.91 |
| 10 | 13,627 | 1,310 | €444.48 | €433.00 | 1.03 |
Cohorts require at least 500 final policies. Empirical intervals omit unseen catastrophic losses and unobserved dependence. The capped analysis changes the target; it does not establish the full-loss model’s performance.
Per-claim cap sensitivity and geographic transfer
€50,000 per-claim sensitivity
The capped-target deviance difference is -0.721 (95% interval -1.111 to -0.317), favoring boosting. All models are refitted against capped outcomes using their uncapped development-selected settings. Capped observed loss is €130.75 per policy-year; boosting predicts €123.93. This is a separate estimand, not a contractual coverage limit.
Reserved Ile-de-France policies
All 69,789 policies and 2,591 claims in this region are withheld from geographic-stress fitting and selection. Both models omit Region as a predictor and fit 485,915 non-reserved policies. The paired deviance difference is -4.567 (95% interval -6.017 to -2.791).
| Model | Deviance | Predicted / observed loss | 95% ratio interval |
|---|---|---|---|
| Poisson–Gamma regression | 83.010 | 1.842 | 1.615–2.123 |
| Boosted frequency–severity | 78.443 | 1.191 | 1.045–1.374 |
The better relative score still accompanies overprediction. A new-region test cannot establish future-time performance.
Tail concentration, segment gaps and count dispersion
Across the source, 88 claims above €50,000 account for €17.92 million of loss; the single largest claim is €4.08 million, or 6.8% of all loss. Capping changes the evidence materially.
The boosted 60+ driver cohort has a predicted/observed ratio of 0.721 (95% interval 0.546–1.005), on 22,637 policies and 887 claims. Area F has ratio 0.604 (0.315–1.311), on 3,628 policies and 150 claims. These descriptive slices have wide uncertainty and do not establish causal effects or fairness.
Mean squared Pearson count residuals are 1.778 for Poisson regression and 1.741 for boosting, above the Poisson reference of 1. The analysis uses empirical policy bootstrap intervals; it does not claim Poisson claim-count prediction intervals.
Paired policy bootstrap: 500 final and geographic-stress draws, 250 per segment; 95% percentile intervals conditional on fitted models and the observed empirical loss tail. No independent-year or catastrophe uncertainty claim. The metrics use 500 draws and the explorer’s segment tables use 250; their portfolio-ratio intervals can differ slightly from Monte Carlo variation.
04 / DATA
Reconcile every claim to its exposure.
The archived CASdatasets 1.2-0 files contain 677,991 policies and 26,444 claims from an unidentified French motor insurer. Every claim joins to a policy and every policy’s recorded claim count agrees with its severity rows. There are no missing source fields.
The analysis retains 235 identical severity rows because equal settlement amounts can represent distinct claims; deleting them breaks count reconciliation. It also retains 1,224 exposures longer than one year, up to 2.01 years. Total exposure is 358,482.8 policy-years and total nominal claim loss is €59.91 million.
Policy IDs join the files and allocate splits. They never enter the model. Exposure is a denominator and weight; outcomes never enter the predictors. R category labels are preserved so internal factor codes cannot corrupt policy joins. Raw policies and individual predictions are not published.
The approximate source period is 2011–2013. There are no reliable policy dates, independently verified feature timestamps or claim-development histories. Some claim amounts reflect French IRSA/IDA settlement conventions. No temporal validation, ultimate-loss estimate or inflation adjustment is claimed.
Source: Dutang, Charpentier & Gallic, Insurance dataset, Recherche Data Gouv, version 1.1, July 12, 2024; CASdatasets 1.2-0 archive. Publisher dictionary. Dataset metadata uses Etalab Open Licence 2.0; package code is separately licensed GPL ≥2. Original aggregate analysis is presented with attribution; no publisher endorsement is implied.
05 / METHOD & LIMITATIONS
Separate model selection, full-loss evaluation and stress tests.
A fixed salted policy-ID hash allocates 405,817 development policies, 135,903 validation policies and 136,271 final policies. Model settings are selected on validation deviance, then refitted on 541,720 development/validation policies. The final cohort is never used for tuning or recalibration.
Poisson frequency uses exposure weights, equivalent to a log-exposure offset. Gamma severity uses claim-count weights. The interpretable models use training-fitted age splines, transformed density and bonus-malus, and risk categories. Direct Tweedie estimates the full pure-premium mean. Histogram boosting uses Poisson frequency and Gamma severity with the same exposure and claim-count identities.
Validation selects regularization .0001 for both regression families and seven leaves for boosting. All candidate comparisons remain in the result artifact. Bootstrap intervals condition on the fitted models and observed policies; shared drivers, catastrophic shocks and future market change remain unresolved.
- Historical French motor data do not validate a current insurance price or rate filing.
- No policy dates support a temporal holdout or ultimate-claim development model.
- Nominal euros are not inflation-adjusted and some amounts reflect settlement conventions.
- Heavy losses and unobserved dependence can make empirical bootstrap intervals too narrow for future risk.
- Driver age and geography are descriptive benchmark features, not causal effects or a fairness certification.
- No policyholder eligibility or pricing action is taken; expense scenarios exclude taxes, profit, capital and regulatory constraints.
- Independent technical review is pending.
06 / CODE
Reproduce the loss and exposure accounting.
Run S31-02c309de-1e141502
Analysis commit 02c309de271f6bb119a598b0c5972a3a697e553b
git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S31/study.py
uv run python scripts/report_s31.py
uv run python -W error -m unittest discover -s tests -vPinned publisher archive and extracted-file hashes, locked dependencies and fixed seeds. Checks cover policy joins, factor labels, exposure and severity identities, split stability, tail definitions, uncertainty and display accounting. Independent technical review is pending.