01 / DECISION · EXECUTIVE SUMMARY
Improve the forecast, then choose how much service is worth.
A merchandise planner needs to balance shelf availability against excess stock. In this bounded retail study, a model trained across products and stores forecasts the next 28-day sales total more accurately than repeating the previous cycle. That is a useful improvement, but it does not choose the business’s service target.
What a leader should do next
Validate the challenger on a current assortment with actual stock availability and costs. Agree on the service target before choosing the order policy. In the default scenario, the challenger costs $7,775.46 versus $7,962.33 for a conventional target, while missing 4,190 versus 4,149 scenario units.
What this does not establish
These are historical sales forecasts and hypothetical inventory replays. The data does not reveal unconstrained customer demand, real inventory or retailer margins. The dollar comparison is not realized savings, and the 21 selected items do not represent every retailer. Total-assortment forecast bands cover only two of the four final windows, so aggregate uncertainty still needs improvement.
02 / IMPLICATION
Forecast quality and operating policy answer different questions.
The challenger’s primary error is 0.8049, versus 0.9746 for repeating the last cycle. The paired difference is -0.1697, with a 95% interval from -0.2554 to -0.0885. Negative values favor the challenger.
In the default replay, lower holding cost more than offsets a small increase in unmet units. A business with a stricter service commitment may prefer a different storage or shortage-cost assumption. The controls show that tradeoff explicitly.
03 / EVIDENCE
Measured forecasts. Clearly labeled scenario economics.
Final evidence covers 30,600 recorded sales units across 840 product/store cycles. The primary metric is weighted, scaled absolute error of a 28-day total; lower is better. It is not the competition’s daily leaderboard score.
| Model | Cycle error | 95% interval | Scaled pinball loss | 80% band coverage |
|---|---|---|---|---|
| Repeat the last week | 1.2265 | 1.0864 to 1.4269 | 1.3226 | 74.6% |
| Repeat the last cycle | 0.9746 | 0.8079 to 1.1151 | 1.2186 | 76.4% |
| Global quantile boosting | 0.8049 | 0.6935 to 0.8982 | 1.1187 | 88.3% |
Intervals resample 21 item clusters with their stores and dates held together. They are conditional on this fixed assortment, historical period and fitted models.
FORECAST & INVENTORY EXPLORER
Choose the service–cost tradeoff explicitly.
Forecasts are saved historical evaluations. Economic controls select one of 36 precomputed simulations; they do not change measured forecast accuracy or retrain the models.
| Origin | Observed units | Forecast units | 80% range | Covered? |
|---|---|---|---|---|
| 2016-02-28 | 7,807 | 7,293 | 4,815–7,560 | No |
| 2016-03-27 | 7,589 | 7,353 | 4,927–7,796 | Yes |
| 2016-04-24 | 7,454 | 7,063 | 4,773–7,497 | Yes |
| 2016-05-22 | 7,750 | 7,143 | 4,812–6,540 | No |
Bands use 17 overlapping calibration histories and are not guaranteed to cover future outcomes. Summing bottom-level scenarios preserves accounting; aggregate quantiles are calculated after summing. The summed point forecast is a separate statistic and can fall outside the aggregate scenario band.
ASSUMED ECONOMICS / SCENARIO ONLY
One order, limited space and an explicit arrival date.
Observed daily sales are replayed as hypothetical demand. Actual inventory availability and costs are unknown. All results below combine the same 210 series and four final windows; the forecast aggregation selector changes only the chart above.
Cost difference versus the conventional target: -$186.87; paired 95% item-cluster interval -$499.28 to -$18.61. Negative values favor the selected policy under these assumptions.
| Storage multiplier | Assumed cost | Fill rate |
|---|---|---|
| 0.5× | $16,056.54 | 50.7% |
| 1× | $7,775.46 | 86.3% |
| 2× | $7,058.35 | 93.0% |
| Policy | Assumed cost | Lost units | Hindsight regret |
|---|---|---|---|
| Repeat the last week | $8,274.43 | 4,761 | $1,492.21 |
| Repeat the last cycle | $7,958.71 | 4,151 | $1,176.49 |
| Global quantile boosting | $7,775.46 | 4,190 | $993.24 |
| Conventional prior-cycle target | $7,962.33 | 4,149 | $1,180.11 |
Order quantities and scenario accounting
| Source-day origin | Order units | Fill rate | Assumed cost |
|---|---|---|---|
| d_1857 | 6,995 | 87.4% | $1,922.12 |
| d_1885 | 6,654 | 85.6% | $2,005.20 |
| d_1913 | 6,399 | 87.8% | $1,790.16 |
| d_1941 | 6,403 | 84.6% | $2,057.98 |
Initial stock equals prior average daily sales × lead time, rounded up and capped by space. The single order arrives before demand on day 4. Space for the entire order is reserved at the origin. Lost demand is not backordered.
Forecast order target: quantile at penalty ÷ (penalty + 28 × daily holding cost), minus initial stock and capped by available space. Cost = end-of-day stock × holding cost + lost units × penalty. No purchase cost or actual profit is measured.
Hindsight chooses the exact best feasible integer order under the same constraints: $6,782.22. It knows future scenario demand, so its cost is a comparator, not an available operational forecast. Uncertainty resamples fixed item clusters, not future market conditions.
Where performance changes across stores and dates
| Store | Origin day | Cycle error | 80% band coverage |
|---|---|---|---|
| CA_1 | d_1857 | 1.049 | 85.7% |
| CA_2 | d_1857 | 0.909 | 81.0% |
| CA_3 | d_1857 | 0.828 | 95.2% |
| CA_4 | d_1857 | 0.928 | 95.2% |
| TX_1 | d_1857 | 1.359 | 76.2% |
| TX_2 | d_1857 | 0.787 | 90.5% |
| TX_3 | d_1857 | 0.656 | 95.2% |
| WI_1 | d_1857 | 0.661 | 85.7% |
| WI_2 | d_1857 | 0.361 | 95.2% |
| WI_3 | d_1857 | 0.856 | 85.7% |
| CA_1 | d_1885 | 0.683 | 95.2% |
| CA_2 | d_1885 | 1.121 | 76.2% |
| CA_3 | d_1885 | 0.562 | 95.2% |
| CA_4 | d_1885 | 0.389 | 90.5% |
| TX_1 | d_1885 | 0.730 | 90.5% |
| TX_2 | d_1885 | 0.865 | 95.2% |
| TX_3 | d_1885 | 0.470 | 85.7% |
| WI_1 | d_1885 | 0.937 | 90.5% |
| WI_2 | d_1885 | 1.608 | 81.0% |
| WI_3 | d_1885 | 0.446 | 90.5% |
| CA_1 | d_1913 | 0.535 | 95.2% |
| CA_2 | d_1913 | 1.254 | 85.7% |
| CA_3 | d_1913 | 0.603 | 85.7% |
| CA_4 | d_1913 | 0.626 | 100.0% |
| TX_1 | d_1913 | 0.900 | 81.0% |
| TX_2 | d_1913 | 0.999 | 95.2% |
| TX_3 | d_1913 | 0.746 | 81.0% |
| WI_1 | d_1913 | 0.624 | 90.5% |
| WI_2 | d_1913 | 0.769 | 90.5% |
| WI_3 | d_1913 | 0.879 | 90.5% |
| CA_1 | d_1941 | 0.788 | 95.2% |
| CA_2 | d_1941 | 0.699 | 85.7% |
| CA_3 | d_1941 | 0.737 | 85.7% |
| CA_4 | d_1941 | 0.751 | 90.5% |
| TX_1 | d_1941 | 1.152 | 85.7% |
| TX_2 | d_1941 | 0.634 | 90.5% |
| TX_3 | d_1941 | 0.696 | 85.7% |
| WI_1 | d_1941 | 1.133 | 71.4% |
| WI_2 | d_1941 | 0.877 | 81.0% |
| WI_3 | d_1941 | 0.522 | 90.5% |
All department/category failures, aggregate levels, calibration origins and tuning scores are in the downloadable JSON and full report.
How dependence assumptions change total forecast bands
| Origin day | Scenario dependence | Band width in units | Covered? |
|---|---|---|---|
| d_1857 | Shared historical ranks | 2,745 | No |
| d_1857 | Independent ranks | 1,358 | No |
| d_1885 | Shared historical ranks | 2,869 | Yes |
| d_1885 | Independent ranks | 1,523 | No |
| d_1913 | Shared historical ranks | 2,723 | Yes |
| d_1913 | Independent ranks | 1,461 | No |
| d_1941 | Shared historical ranks | 1,728 | No |
| d_1941 | Independent ranks | 1,389 | No |
Seventeen overlapping histories offer limited information about joint tails. The comparison tests a dependence assumption; neither method guarantees future coverage.
04 / DATA
Real unit sales, a fixed assortment and a conservative price cutoff.
The M5 organizer’s sales files cover January 29, 2011 through June 19, 2016. The source has 30,490 product/store series and 6.84 million weekly price rows. The study selects three eligible items per department by a fixed identifier hash, retaining all ten stores: 21 items and 210 series.
Selection uses only identifiers and price availability before tuning. Training and evaluation files join on item/store; the calendar is consecutive and all selected final origins have eligible prices. About 71.4% of the selected pre-final daily observations are zero sales. Zeros are recorded sales, not proof of absent demand.
Only completed-week prices and past sales enter a forecast. No future promotion outcome or price is supplied. Dates in the final chart are the forecast cutoff, followed by 28 days of observed sales.
Official M5 source and access documentation ↗
Makridakis, Spiliotis & Assimakopoulos (2022), M5 accuracy competition; official M5-methods repository and competitors guide. Raw files are not redistributed. Exact file hashes and provenance are recorded in the result artifact.
05 / METHOD & LIMITATIONS
Rolling origins with fixed model settings.
A single earlier tuning origin chooses between 7 and 15 leaves using quantile loss. The selected 15-leaf models predict the 10th, 50th and 90th percentiles. Seventeen subsequent calibration origins inform forecast scenarios. Every training outcome must finish before its forecast origin.
The four final 28-day windows do not overlap. Models refit on eligible history at each origin, without retuning. Whole historical residual-rank vectors preserve cross-series dependence in scenarios. Every store/category/total draw is the sum of bottom-level units; quantiles are taken after aggregation.
Inventory replay makes one constrained order, with explicit lead time, initial stock and reserved capacity. A tested discrete convex search finds the exact hindsight order under the same constraints. Its advantage is information about the future, not evidence of an implementable forecast.
- Bounded deterministic 21-item assortment; not a random sample of the full retail market.
- Cycle-total scaled error is not the official daily M5 leaderboard score.
- Only four final historical windows; item-bootstrap intervals do not quantify future time-regime uncertainty.
- Seventeen overlapping calibration windows limit tail and dependence estimation; forecast bands have no exact-coverage guarantee.
- No actual stock availability, retailer margin, purchase cost, spoilage or realized savings is observed.
- Independent technical review is pending.
06 / CODE
Reproduce the comparison from the recorded files.
Run S02-5145fea6-e598f30c
Analysis commit 5145fea621938bb3455655a685897dfa69a2cd83
git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
# Download the four official files listed in studies/S02/DATA.md.
RESEARCH_DATA_DIR=/absolute/path/to/data uv run python -W error studies/S02/study.py
uv run python scripts/report_s02.py
uv run python -W error -m unittest discover -s tests -vPython 3.11.16 · locked dependencies · fixed seeds. Focused checks cover future-data leakage, stock conservation, exact hindsight optimization, scenario coherence and saved result accounting. Independent technical review is pending.