← Research catalog

S13 / MANUFACTURING

More sensor data does not guarantee dependable screening.

The tested sensor models provide weak later-month ranking evidence. At a 20% inspection budget, boosting identifies only 4 of 22 failures; its comparison with sparse logistic is inconclusive.

Evaluated public-data studyResearch code ↗Executive summary ↓

359 final production entities · 22 failures · October 1–17, 2008 · SECOM

01 / DECISION · EXECUTIVE SUMMARY

Keep the quality gate until stronger evidence supports changing it.

A manufacturing leader needs a screening system that finds failures without overwhelming inspection capacity. In this historical comparison, hundreds of sensor signals do not produce a dependable priority queue in the later month.

4 / 22failures found by boosting
71 / 359units inspected at 20% capacity
18failures left outside that queue

The model upgrade is not established.

The sparse model identifies 3 failures at the same budget; the PCA monitor identifies 5. Across all ranking thresholds, the boosting-versus-sparse average-precision difference is -0.0140, with a 95% interval of -0.0963 to +0.0212. The evidence does not support a clear advantage.

This is not a validated early-warning system.

Sensor acquisition times are not supplied. The comparison assumes the recorded measurements would be available at screening. It cannot establish advance warning, engineering root causes, avoided failures or savings.

02 / IMPLICATION

Require performance that survives the next production period.

Boosting’s validation average precision was 0.2519 in September; it falls to 0.0626 in October, close to the final failure prevalence of 0.0613. This is a reason to demand prospective validation and verified information timing before changing an operational quality gate.

A constant training-prevalence forecast has better final log loss and Brier score than the fitted sensor models. Adding complexity did not produce dependable probabilities here. The study has not tested an inspection intervention or a current manufacturing process.

03 / EVIDENCE

Use the later month, not the most flattering split.

The final cohort contains 359 entities, 22 failures and 17 calendar days. Average precision summarizes precision across recall thresholds; higher is better. It is not an accuracy percentage or the failure-capture rate at a particular inspection budget.

October final evaluation · paired day-cluster uncertainty
ModelAverage precision95% intervalLog lossBrier score
Sparse elastic-net logistic0.07650.0345–0.17480.25500.0602
PCA process monitor0.06600.0318–0.13350.23580.0585
Gradient boosting0.06260.0317–0.09930.25860.0594
Boosting without missingness indicators0.06320.0312–0.10200.25860.0594
Training failure prevalence0.06130.0252–0.09770.23080.0576

HISTORICAL INSPECTION EXPLORER

How many failures remain outside the review queue?

Controls display saved October results. The inspection share is an assumed batch budget; moving it does not refit a model or prevent a failure.

4 / 22observed failures identified
18observed failures outside the queue
67inspected units that passed

18.2% of failures captured. Queue precision: 5.6%. These are exact counts in one historical cohort; future detection is uncertain.

This model’s average precision across thresholds is 0.0626 (95% day-cluster interval 0.0317–0.0993). That metric differs from recall at the chosen budget. Only 17 final days and 22 failures support the evaluation.

Observed failures identified as batch inspection capacity increases051015200%25%50%75%100%Failures identified · selected model
Horizontal: assumed share of the complete final batch inspected. Vertical: observed failures captured. At 100%, all 22 labels are recovered by construction. The curve is a historical ranking replay, not measured defect reduction.
20% inspection budget · same 359 final entities
ModelInspectionsDetected failuresMissed failuresInspected passes
Sparse elastic-net logistic7131968
PCA process monitor7151766
Gradient boosting7141867
Boosting without missingness indicators7122069
Probability calibration for the selected model
Five equal-count score groups · descriptive, not recalibrated
GroupUnitsFailuresMean predicted riskObserved failure rate
17251.2%6.9%
27241.8%5.6%
37232.4%4.2%
47263.3%8.3%
57146.5%5.6%

Which anonymous signals are repeatedly selected?

These results always describe the sparse logistic model’s 30 training-day bootstrap refits. They are separate from the selected screening model above. Frequency is not physical importance or proof of a fault.

10 sensors meet this display filter. Showing the first 10 by selection frequency; the aggregate download contains all recorded sensors.

Anonymous source columns · 30 refits
SensorSelectedAvailable after filteringIn final fit
sensor_060100.0%100.0%Yes
sensor_02296.7%100.0%Yes
sensor_46193.3%100.0%Yes
sensor_33490.0%100.0%Yes
sensor_56390.0%100.0%Yes
sensor_06586.7%100.0%Yes
sensor_00183.3%100.0%Yes
sensor_01583.3%100.0%Yes
sensor_10383.3%100.0%Yes
sensor_48783.3%100.0%Yes

Model uncertainty conditions on the fitted models and this historical period. Sensors lack physical names and acquisition timestamps. A display threshold changes the table only; it does not create a new sensor model or alter held-out accuracy.

Selection stability, missingness and dependence sensitivity

The final sparse model uses 79 nonzero coefficients across 78 sensors. The 30 training-day resamples select 72–129 sensors, with mean Jaccard overlap of 32.5% against the final set. Even a frequently selected signal is not an identified process fault.

Removing missingness indicators leaves boosting average precision near 0.0632 and captures two failures at 20% capacity. Forty-eight final entities have more than 5% missing sensors and no observed failures; that slice cannot validate failure detection. The fixed two-day block sensitivity for the primary difference is -0.0676 to +0.0145, also inconclusive.

1,000 paired calendar-day cluster bootstrap draws; 95% percentile intervals condition on the fitted models and 17 observed test days. Two-day circular-block sensitivity is supplementary. Sensor stability uses 30 training-day bootstrap refits.

04 / DATA

Inspect the measurements and the calendar first.

SECOM supplies 1,567 production entities and 104 failures from July 19–October 17, 2008. The numeric file contains 590 anonymous sensor columns, although the documentation lists 591. Labels and quoted test timestamps are aligned in a separate file.

The source has 41,951 missing sensor cells, 116 constant columns, 104 exact duplicate columns and 33 timestamp ties. There are no exact duplicate complete sensor rows. All removal and transformation rules are learned again inside each permitted training fit.

July–August supplies 618 development entities and 65 failures; September supplies 590 validation entities and 17 failures. The final selected models are refitted on 1,208 July–September entities, then evaluated once on October. Timestamps order the split and never enter a predictor. Unknown label delays and unobserved production batches limit the retrospective design.

Source: McCann & Johnston (2008), SECOM, UCI Machine Learning Repository, doi:10.24432/C54305, CC BY 4.0. The study transforms those files into aggregate evaluations and original figures; no publisher endorsement is implied.

05 / METHOD & LIMITATIONS

Separate preprocessing, model selection and final evidence.

Training-only missingness and variation filters, duplicate removal, clipping, median imputation and standardization produce the model inputs. Elastic-net logistic, a PCA process monitor and gradient boosting each have a small, frozen setting grid selected by September log loss.

The PCA monitor learns normal sensor variation from training passes and maps its residual and component scores to failure probability using training labels. The selected logistic model uses C=.1; the selected PCA monitor has five components; boosting has seven leaves. A fixed no-missingness ablation and 30 day-cluster refits probe fragile signals. No extra calibration is fitted on the small validation cohort.

  • Only 22 final failures and 17 final days support limited precision.
  • Anonymous sensors and missing acquisition times prevent root-cause and validated early-warning claims.
  • A short historical sample with changing prevalence does not validate a modern production line.
  • Intervals condition on fitted models and cannot account for unrecorded batch dependence.
  • Selection stability is descriptive and does not establish causal or physical sensor importance.
  • Inspection counts are historical ranking outcomes, not measured savings or avoided failures.
  • Independent technical review is pending.
Frozen calendar and sensor protocol ↗

06 / CODE

Trace the evidence to the recorded run.

Run S13-608719b7-eea568ba
Analysis commit 608719b77b23e68f298f456d69b725aba42a8d16

git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S13/study.py
uv run python scripts/report_s13.py
uv run python -W error -m unittest discover -s tests -v

Pinned source hashes, locked dependencies and fixed seeds. Checks cover calendar boundaries, training-only preprocessing, probability optimization, cluster resampling, capacity accounting and display-to-result agreement. Individual records remain outside publication. Independent technical review is pending.