← Research catalog

S02 / RETAIL AND E-COMMERCE

Better forecasts still need a service decision.

Quantile boosting reduces cycle forecast error by 17.4% in the tested assortment. Lower simulated cost does not always mean fewer missed sales.

Evaluated public-data studyResearch code ↗Executive summary ↓

210 product/store series · 10 stores · 4 final 28-day windows · M5 sales through June 2016

01 / DECISION · EXECUTIVE SUMMARY

Improve the forecast, then choose how much service is worth.

A merchandise planner needs to balance shelf availability against excess stock. In this bounded retail study, a model trained across products and stores forecasts the next 28-day sales total more accurately than repeating the previous cycle. That is a useful improvement, but it does not choose the business’s service target.

17.4%lower primary forecast error
840final product/store cycles
36explicit inventory scenarios

What a leader should do next

Validate the challenger on a current assortment with actual stock availability and costs. Agree on the service target before choosing the order policy. In the default scenario, the challenger costs $7,775.46 versus $7,962.33 for a conventional target, while missing 4,190 versus 4,149 scenario units.

What this does not establish

These are historical sales forecasts and hypothetical inventory replays. The data does not reveal unconstrained customer demand, real inventory or retailer margins. The dollar comparison is not realized savings, and the 21 selected items do not represent every retailer. Total-assortment forecast bands cover only two of the four final windows, so aggregate uncertainty still needs improvement.

02 / IMPLICATION

Forecast quality and operating policy answer different questions.

The challenger’s primary error is 0.8049, versus 0.9746 for repeating the last cycle. The paired difference is -0.1697, with a 95% interval from -0.2554 to -0.0885. Negative values favor the challenger.

In the default replay, lower holding cost more than offsets a small increase in unmet units. A business with a stricter service commitment may prefer a different storage or shortage-cost assumption. The controls show that tradeoff explicitly.

03 / EVIDENCE

Measured forecasts. Clearly labeled scenario economics.

Final evidence covers 30,600 recorded sales units across 840 product/store cycles. The primary metric is weighted, scaled absolute error of a 28-day total; lower is better. It is not the competition’s daily leaderboard score.

All prespecified models · final four windows
ModelCycle error95% intervalScaled pinball loss80% band coverage
Repeat the last week1.22651.0864 to 1.42691.322674.6%
Repeat the last cycle0.97460.8079 to 1.11511.218676.4%
Global quantile boosting0.80490.6935 to 0.89821.118788.3%

Intervals resample 21 item clusters with their stores and dates held together. They are conditional on this fixed assortment, historical period and fitted models.

FORECAST & INVENTORY EXPLORER

Choose the service–cost tradeoff explicitly.

Forecasts are saved historical evaluations. Economic controls select one of 36 precomputed simulations; they do not change measured forecast accuracy or retrain the models.

Observed sales and saved 80% forecast bandsFour final 28-day windows for the selected model and group. Exact quantities and coverage are in the following table.01,9523,9045,8557,807Window 1Window 2Window 3Window 4Recorded sales units over 28 days
Slate bands: empirical 80% forecast ranges. Copper dots: summed point forecasts. Navy crosses: observed sales. Dates below are forecast origins. On narrow screens, scroll the chart horizontally.
selected assortment · historical forecasts
OriginObserved unitsForecast units80% rangeCovered?
2016-02-287,8077,2934,815–7,560No
2016-03-277,5897,3534,927–7,796Yes
2016-04-247,4547,0634,773–7,497Yes
2016-05-227,7507,1434,812–6,540No

Bands use 17 overlapping calibration histories and are not guaranteed to cover future outcomes. Summing bottom-level scenarios preserves accounting; aggregate quantiles are calculated after summing. The summed point forecast is a separate statistic and can fall outside the aggregate scenario band.

ASSUMED ECONOMICS / SCENARIO ONLY

One order, limited space and an explicit arrival date.

Observed daily sales are replayed as hypothetical demand. Actual inventory availability and costs are unknown. All results below combine the same 210 series and four final windows; the forecast aggregation selector changes only the chart above.

$7,775.46total assumed cost
86.3%scenario unit fill rate
26,451ordered units across four cycles

Cost difference versus the conventional target: -$186.87; paired 95% item-cluster interval -$499.28 to -$18.61. Negative values favor the selected policy under these assumptions.

Storage capacity, simulated service and assumed costThree points show one-half, one and two times prior-cycle sales as storage capacity. Table values follow.$00%$4,01425%$8,02850%$12,04275%$16,057100%0.5×1×2×Assumed total cost (USD)Scenario unit fill rate (%)
Saved sensitivity points for the selected forecast and costs. Larger point: chosen storage. This finite grid is not a globally optimized frontier.
Selected policy · storage sensitivity
Storage multiplierAssumed costFill rate
0.5×$16,056.5450.7%
1×$7,775.4686.3%
2×$7,058.3593.0%
All policies · identical scenario constraints
PolicyAssumed costLost unitsHindsight regret
Repeat the last week$8,274.434,761$1,492.21
Repeat the last cycle$7,958.714,151$1,176.49
Global quantile boosting$7,775.464,190$993.24
Conventional prior-cycle target$7,962.334,149$1,180.11
Order quantities and scenario accounting
Selected policy · aggregate order quantities
Source-day originOrder unitsFill rateAssumed cost
d_18576,99587.4%$1,922.12
d_18856,65485.6%$2,005.20
d_19136,39987.8%$1,790.16
d_19416,40384.6%$2,057.98

Initial stock equals prior average daily sales × lead time, rounded up and capped by space. The single order arrives before demand on day 4. Space for the entire order is reserved at the origin. Lost demand is not backordered.

Forecast order target: quantile at penalty ÷ (penalty + 28 × daily holding cost), minus initial stock and capped by available space. Cost = end-of-day stock × holding cost + lost units × penalty. No purchase cost or actual profit is measured.

Hindsight chooses the exact best feasible integer order under the same constraints: $6,782.22. It knows future scenario demand, so its cost is a comparator, not an available operational forecast. Uncertainty resamples fixed item clusters, not future market conditions.

Where performance changes across stores and dates
Challenger · all stores and final origins
StoreOrigin dayCycle error80% band coverage
CA_1d_18571.04985.7%
CA_2d_18570.90981.0%
CA_3d_18570.82895.2%
CA_4d_18570.92895.2%
TX_1d_18571.35976.2%
TX_2d_18570.78790.5%
TX_3d_18570.65695.2%
WI_1d_18570.66185.7%
WI_2d_18570.36195.2%
WI_3d_18570.85685.7%
CA_1d_18850.68395.2%
CA_2d_18851.12176.2%
CA_3d_18850.56295.2%
CA_4d_18850.38990.5%
TX_1d_18850.73090.5%
TX_2d_18850.86595.2%
TX_3d_18850.47085.7%
WI_1d_18850.93790.5%
WI_2d_18851.60881.0%
WI_3d_18850.44690.5%
CA_1d_19130.53595.2%
CA_2d_19131.25485.7%
CA_3d_19130.60385.7%
CA_4d_19130.626100.0%
TX_1d_19130.90081.0%
TX_2d_19130.99995.2%
TX_3d_19130.74681.0%
WI_1d_19130.62490.5%
WI_2d_19130.76990.5%
WI_3d_19130.87990.5%
CA_1d_19410.78895.2%
CA_2d_19410.69985.7%
CA_3d_19410.73785.7%
CA_4d_19410.75190.5%
TX_1d_19411.15285.7%
TX_2d_19410.63490.5%
TX_3d_19410.69685.7%
WI_1d_19411.13371.4%
WI_2d_19410.87781.0%
WI_3d_19410.52290.5%

All department/category failures, aggregate levels, calibration origins and tuning scores are in the downloadable JSON and full report.

How dependence assumptions change total forecast bands
Challenger · selected-assortment total
Origin dayScenario dependenceBand width in unitsCovered?
d_1857Shared historical ranks2,745No
d_1857Independent ranks1,358No
d_1885Shared historical ranks2,869Yes
d_1885Independent ranks1,523No
d_1913Shared historical ranks2,723Yes
d_1913Independent ranks1,461No
d_1941Shared historical ranks1,728No
d_1941Independent ranks1,389No

Seventeen overlapping histories offer limited information about joint tails. The comparison tests a dependence assumption; neither method guarantees future coverage.

04 / DATA

Real unit sales, a fixed assortment and a conservative price cutoff.

The M5 organizer’s sales files cover January 29, 2011 through June 19, 2016. The source has 30,490 product/store series and 6.84 million weekly price rows. The study selects three eligible items per department by a fixed identifier hash, retaining all ten stores: 21 items and 210 series.

Selection uses only identifiers and price availability before tuning. Training and evaluation files join on item/store; the calendar is consecutive and all selected final origins have eligible prices. About 71.4% of the selected pre-final daily observations are zero sales. Zeros are recorded sales, not proof of absent demand.

Only completed-week prices and past sales enter a forecast. No future promotion outcome or price is supplied. Dates in the final chart are the forecast cutoff, followed by 28 days of observed sales.

Official M5 source and access documentation ↗

Makridakis, Spiliotis & Assimakopoulos (2022), M5 accuracy competition; official M5-methods repository and competitors guide. Raw files are not redistributed. Exact file hashes and provenance are recorded in the result artifact.

05 / METHOD & LIMITATIONS

Rolling origins with fixed model settings.

A single earlier tuning origin chooses between 7 and 15 leaves using quantile loss. The selected 15-leaf models predict the 10th, 50th and 90th percentiles. Seventeen subsequent calibration origins inform forecast scenarios. Every training outcome must finish before its forecast origin.

The four final 28-day windows do not overlap. Models refit on eligible history at each origin, without retuning. Whole historical residual-rank vectors preserve cross-series dependence in scenarios. Every store/category/total draw is the sum of bottom-level units; quantiles are taken after aggregation.

Inventory replay makes one constrained order, with explicit lead time, initial stock and reserved capacity. A tested discrete convex search finds the exact hindsight order under the same constraints. Its advantage is information about the future, not evidence of an implementable forecast.

  • Bounded deterministic 21-item assortment; not a random sample of the full retail market.
  • Cycle-total scaled error is not the official daily M5 leaderboard score.
  • Only four final historical windows; item-bootstrap intervals do not quantify future time-regime uncertainty.
  • Seventeen overlapping calibration windows limit tail and dependence estimation; forecast bands have no exact-coverage guarantee.
  • No actual stock availability, retailer margin, purchase cost, spoilage or realized savings is observed.
  • Independent technical review is pending.

Frozen protocol and split definitions ↗

06 / CODE

Reproduce the comparison from the recorded files.

Run S02-5145fea6-e598f30c
Analysis commit 5145fea621938bb3455655a685897dfa69a2cd83

git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
# Download the four official files listed in studies/S02/DATA.md.
RESEARCH_DATA_DIR=/absolute/path/to/data uv run python -W error studies/S02/study.py
uv run python scripts/report_s02.py
uv run python -W error -m unittest discover -s tests -v

Python 3.11.16 · locked dependencies · fixed seeds. Focused checks cover future-data leakage, stock conservation, exact hindsight optimization, scenario coherence and saved result accounting. Independent technical review is pending.