01 / DECISION · EXECUTIVE SUMMARY
Keep the quality gate until stronger evidence supports changing it.
A manufacturing leader needs a screening system that finds failures without overwhelming inspection capacity. In this historical comparison, hundreds of sensor signals do not produce a dependable priority queue in the later month.
The model upgrade is not established.
The sparse model identifies 3 failures at the same budget; the PCA monitor identifies 5. Across all ranking thresholds, the boosting-versus-sparse average-precision difference is -0.0140, with a 95% interval of -0.0963 to +0.0212. The evidence does not support a clear advantage.
This is not a validated early-warning system.
Sensor acquisition times are not supplied. The comparison assumes the recorded measurements would be available at screening. It cannot establish advance warning, engineering root causes, avoided failures or savings.
02 / IMPLICATION
Require performance that survives the next production period.
Boosting’s validation average precision was 0.2519 in September; it falls to 0.0626 in October, close to the final failure prevalence of 0.0613. This is a reason to demand prospective validation and verified information timing before changing an operational quality gate.
A constant training-prevalence forecast has better final log loss and Brier score than the fitted sensor models. Adding complexity did not produce dependable probabilities here. The study has not tested an inspection intervention or a current manufacturing process.
03 / EVIDENCE
Use the later month, not the most flattering split.
The final cohort contains 359 entities, 22 failures and 17 calendar days. Average precision summarizes precision across recall thresholds; higher is better. It is not an accuracy percentage or the failure-capture rate at a particular inspection budget.
| Model | Average precision | 95% interval | Log loss | Brier score |
|---|---|---|---|---|
| Sparse elastic-net logistic | 0.0765 | 0.0345–0.1748 | 0.2550 | 0.0602 |
| PCA process monitor | 0.0660 | 0.0318–0.1335 | 0.2358 | 0.0585 |
| Gradient boosting | 0.0626 | 0.0317–0.0993 | 0.2586 | 0.0594 |
| Boosting without missingness indicators | 0.0632 | 0.0312–0.1020 | 0.2586 | 0.0594 |
| Training failure prevalence | 0.0613 | 0.0252–0.0977 | 0.2308 | 0.0576 |
HISTORICAL INSPECTION EXPLORER
How many failures remain outside the review queue?
Controls display saved October results. The inspection share is an assumed batch budget; moving it does not refit a model or prevent a failure.
18.2% of failures captured. Queue precision: 5.6%. These are exact counts in one historical cohort; future detection is uncertain.
This model’s average precision across thresholds is 0.0626 (95% day-cluster interval 0.0317–0.0993). That metric differs from recall at the chosen budget. Only 17 final days and 22 failures support the evaluation.
| Model | Inspections | Detected failures | Missed failures | Inspected passes |
|---|---|---|---|---|
| Sparse elastic-net logistic | 71 | 3 | 19 | 68 |
| PCA process monitor | 71 | 5 | 17 | 66 |
| Gradient boosting | 71 | 4 | 18 | 67 |
| Boosting without missingness indicators | 71 | 2 | 20 | 69 |
Probability calibration for the selected model
| Group | Units | Failures | Mean predicted risk | Observed failure rate |
|---|---|---|---|---|
| 1 | 72 | 5 | 1.2% | 6.9% |
| 2 | 72 | 4 | 1.8% | 5.6% |
| 3 | 72 | 3 | 2.4% | 4.2% |
| 4 | 72 | 6 | 3.3% | 8.3% |
| 5 | 71 | 4 | 6.5% | 5.6% |
Which anonymous signals are repeatedly selected?
These results always describe the sparse logistic model’s 30 training-day bootstrap refits. They are separate from the selected screening model above. Frequency is not physical importance or proof of a fault.
10 sensors meet this display filter. Showing the first 10 by selection frequency; the aggregate download contains all recorded sensors.
| Sensor | Selected | Available after filtering | In final fit |
|---|---|---|---|
| sensor_060 | 100.0% | 100.0% | Yes |
| sensor_022 | 96.7% | 100.0% | Yes |
| sensor_461 | 93.3% | 100.0% | Yes |
| sensor_334 | 90.0% | 100.0% | Yes |
| sensor_563 | 90.0% | 100.0% | Yes |
| sensor_065 | 86.7% | 100.0% | Yes |
| sensor_001 | 83.3% | 100.0% | Yes |
| sensor_015 | 83.3% | 100.0% | Yes |
| sensor_103 | 83.3% | 100.0% | Yes |
| sensor_487 | 83.3% | 100.0% | Yes |
Model uncertainty conditions on the fitted models and this historical period. Sensors lack physical names and acquisition timestamps. A display threshold changes the table only; it does not create a new sensor model or alter held-out accuracy.
Selection stability, missingness and dependence sensitivity
The final sparse model uses 79 nonzero coefficients across 78 sensors. The 30 training-day resamples select 72–129 sensors, with mean Jaccard overlap of 32.5% against the final set. Even a frequently selected signal is not an identified process fault.
Removing missingness indicators leaves boosting average precision near 0.0632 and captures two failures at 20% capacity. Forty-eight final entities have more than 5% missing sensors and no observed failures; that slice cannot validate failure detection. The fixed two-day block sensitivity for the primary difference is -0.0676 to +0.0145, also inconclusive.
1,000 paired calendar-day cluster bootstrap draws; 95% percentile intervals condition on the fitted models and 17 observed test days. Two-day circular-block sensitivity is supplementary. Sensor stability uses 30 training-day bootstrap refits.
04 / DATA
Inspect the measurements and the calendar first.
SECOM supplies 1,567 production entities and 104 failures from July 19–October 17, 2008. The numeric file contains 590 anonymous sensor columns, although the documentation lists 591. Labels and quoted test timestamps are aligned in a separate file.
The source has 41,951 missing sensor cells, 116 constant columns, 104 exact duplicate columns and 33 timestamp ties. There are no exact duplicate complete sensor rows. All removal and transformation rules are learned again inside each permitted training fit.
July–August supplies 618 development entities and 65 failures; September supplies 590 validation entities and 17 failures. The final selected models are refitted on 1,208 July–September entities, then evaluated once on October. Timestamps order the split and never enter a predictor. Unknown label delays and unobserved production batches limit the retrospective design.
Source: McCann & Johnston (2008), SECOM, UCI Machine Learning Repository, doi:10.24432/C54305, CC BY 4.0. The study transforms those files into aggregate evaluations and original figures; no publisher endorsement is implied.
05 / METHOD & LIMITATIONS
Separate preprocessing, model selection and final evidence.
Training-only missingness and variation filters, duplicate removal, clipping, median imputation and standardization produce the model inputs. Elastic-net logistic, a PCA process monitor and gradient boosting each have a small, frozen setting grid selected by September log loss.
The PCA monitor learns normal sensor variation from training passes and maps its residual and component scores to failure probability using training labels. The selected logistic model uses C=.1; the selected PCA monitor has five components; boosting has seven leaves. A fixed no-missingness ablation and 30 day-cluster refits probe fragile signals. No extra calibration is fitted on the small validation cohort.
- Only 22 final failures and 17 final days support limited precision.
- Anonymous sensors and missing acquisition times prevent root-cause and validated early-warning claims.
- A short historical sample with changing prevalence does not validate a modern production line.
- Intervals condition on fitted models and cannot account for unrecorded batch dependence.
- Selection stability is descriptive and does not establish causal or physical sensor importance.
- Inspection counts are historical ranking outcomes, not measured savings or avoided failures.
- Independent technical review is pending.
06 / CODE
Trace the evidence to the recorded run.
Run S13-608719b7-eea568ba
Analysis commit 608719b77b23e68f298f456d69b725aba42a8d16
git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S13/study.py
uv run python scripts/report_s13.py
uv run python -W error -m unittest discover -s tests -vPinned source hashes, locked dependencies and fixed seeds. Checks cover calendar boundaries, training-only preprocessing, probability optimization, cluster resampling, capacity accounting and display-to-result agreement. Individual records remain outside publication. Independent technical review is pending.