← Research catalog

S58 / PROFESSIONAL SERVICES AND ENTERPRISE AI

Better delay forecasts. A modest operational signal.

Detailed event history reduces forecasting error by 1.6% compared with knowing an application’s stage and age. Faster actual completion remains an untested intervention.

Evaluated public-data studyResearch code ↗Executive summary ↓

2,365 held-out applications · 6,247 event prefixes · December 2016 cohort · BPI Challenge 2017

01 / DECISION · EXECUTIVE SUMMARY

Keep the simple benchmark when prioritizing slow cases.

An operations leader needs to decide which active applications deserve scarce review time. The detailed workflow model makes slightly better forecasts than a simple stage-and-age model, but the gain is small. It supports testing a prioritization process; it does not justify a claim of faster service or lower staffing cost.

8.72 daysstage/age forecast error
8.58 daysdetailed-history forecast error
1.6%lower mean absolute error

What a leader can take from this

Use stage and age as a credible starting point. At event 20 and 20% review capacity, the detailed model identifies 218 late applications versus 215 for the simple benchmark. Evaluate whether that small targeting difference helps a real review team before investing in a complex deployment.

Forecasting improvement is not process improvement.

The paired error difference is -0.139 days (95% interval -0.181 to -0.093). That means roughly 3.3 hours less forecasting error, not applications completed 3.3 hours sooner. Predictions end at a defined recorded status and cap remaining time at 30 days.

02 / IMPLICATION

Investigate the handoff; test the intervention.

Among common transitions visible in early application prefixes, the gap from A_Complete to A_Validating has a median of 6.86 elapsed days across 266 observations. That suggests a concrete handoff to investigate. The timestamps alone do not distinguish waiting, customer response, batching or active employee work.

A sensible pilot would compare the existing review queue with model prioritization at the same capacity and measure actual downstream outcomes. This historical forecast study does not estimate that intervention. Its staffing controls expose assumptions so visitors can see how strongly a scenario depends on them.

03 / EVIDENCE

Follow the time cutoff, the simple alternatives and the uncertainty.

The final evaluation contains 2,365 distinct applications and 6,247 eligible prefixes. All prefixes from the same application stay together. Each application receives total weight one across its available prefixes.

December applications · all eligible prefix lengths
ModelMAE, days95% interval14-day Brier
Pooled survival baseline9.3269.190–9.4620.2455
Current stage and age baseline8.7228.597–8.8690.2262
Workflow hazard boosting8.5838.459–8.7340.2251
Age and current stage only8.6728.551–8.8150.2245

Lower error and Brier scores are better. The reduced-history hazard model has a slightly better Brier score than the detailed model, so extra history does not win every comparison. The selected model’s nominal 80% prediction bands cover 88.5% of restricted outcomes: conservative coverage, not exact calibration.

WORKFLOW EVIDENCE EXPLORER

Where does a forecast help the review queue?

Explore saved December results. Stage and prefix filters change the evaluated cohort; no control refits a model or simulates new applications.

13.50mean forecast remaining days
13.72mean observed remaining days
7.98mean absolute error, days

1,588 applications · 1,588 prefixes in this selection. Nominal 80% band coverage: 84.9%.

Average individual 80% prediction-band endpoints: 4.01–24.28 days. These averages are neither an interval for the group mean nor a guarantee for a particular application. Times are capped at 30 days. Subsets retain application-prefix weights; cells below 30 prefixes are unavailable.

Observed application-stage transitions

Each card is a recorded link in the process. Select an available destination stage to inspect its forecasts above. The map uses the longest eligible prefix per application, independently of the filters.

Concept →0.22 daysmedian elapsed · 2,217 transitions90th percentile: 1.81 days
Create Application →<0.01 daysmedian elapsed · 1,573 transitions90th percentile: <0.01 days
Submitted →<0.01 daysmedian elapsed · 1,509 transitions90th percentile: 0.03 days
Accepted →<0.01 daysmedian elapsed · 1,320 transitions90th percentile: <0.01 days
Create Application →<0.01 daysmedian elapsed · 792 transitions90th percentile: <0.01 days
Complete →6.86 daysmedian elapsed · 266 transitions90th percentile: 12.77 days

Edges with fewer than 30 observations are omitted. Disabled destinations lack an eligible diagnostic cell at this prefix. Elapsed time is not active work time; early prefixes omit later loops.

Prioritize late applications at a fixed review capacity

This comparison includes all stages at the selected prefix. “Late” means no target status within 14 days after that prefix. Stage filtering above does not change these capacity totals.

218 of 623late applications identified among 317 prioritized applications

Late-case capture: 35.0%; conditional 95% interval 31.1%–38.8%. This measures identification, not improvement caused by review.

Late-case capture by review capacitySelected model and conditional 95% interval compared with the stage and age baseline. Exact values for all models appear in the following table.0%0%25%25%50%50%75%75%100%100%Share of late applications identified (%)Review capacity (%)
Copper: selected model and conditional 95% band. Navy: stage/age baseline. Intervals retain the frozen rankings and are not simultaneous across capacities. Scroll horizontally on narrow screens.
Event 20 · 20% capacity · all eligible stages
ModelLate identifiedCapture95% interval
Pooled survival baseline116 / 62318.6%15.9%–21.6%
Current stage and age baseline215 / 62334.5%31.2%–38.1%
Workflow hazard boosting218 / 62335.0%31.1%–38.8%
Age and current stage only208 / 62333.4%29.9%–37.0%

ASSUMED CAPACITY / SCENARIO ONLY

Separate a staffing assumption from measured evidence.

Assume some share of the selected cohort’s forecast duration varies inversely with capacity. The dataset does not identify that share or the effect of staffing.

13.50 dayshypothetical duration versus 13.50 forecast days

Formula: forecast × [(1 − share) + share ÷ multiplier]. No causal confidence interval is available. Values above 30 days extrapolate beyond the fitted horizon. These assumptions do not alter observed error, measured coverage or the review comparison.

Probability calibration and diagnostic limits
Detailed-history model · probability unresolved after 14 days
Probability binPrefixesPredictedObserved
0–10%2218.0%5.9%
10–20%38814.4%23.7%
20–30%8223.1%42.7%
30–40%20734.1%41.1%
40–50%16846.3%30.3%
50–60%1,23057.3%52.0%
60–70%3,27165.7%64.5%
70–80%68072.2%70.1%

Bins use predictions, without final outcomes. Counts are prefixes, not independent applications. Stage and prefix diagnostics retain the application weights and are withheld below 30 prefixes. Full model comparisons are available in the aggregate download.

500 paired application-cluster bootstrap draws retain all evaluated prefixes of each application. Conditional on fitted models and one institutional/time cohort; no retraining uncertainty. Capacity intervals hold the selected ranking fixed. These are not simultaneous guarantees across every subgroup or slider setting.

04 / DATA

A real workflow with a carefully defined endpoint.

The original BPI Challenge 2017 file contains 31,509 applications and 1,202,267 events. Applications opened in 2016; the last recorded timestamp is February 1, 2017. Offers remain attached to their application, and start, suspend, resume and complete lifecycle events remain distinct.

The target is the first recorded A_Pending, A_Denied or A_Cancelled status. It is not loan disbursement or the last event in the trace. At each of the first 5, 10 or 20 events, features use only information already observed. Resource identities, applicant identifiers and financial attributes do not enter the model or the public tables.

Development uses January–August arrivals; September supplies validation. Final training uses January–September arrivals with eligible prefixes before November and follow-up through November 30. The untouched final cohort consists of December arrivals with prefixes before January 1. That cutoff provides at least 30 days of final follow-up.

Of the final prefixes, 183 lack a recorded endpoint by the source cutoff. Their unrestricted durations are censored, but their 30-day restricted outcome is fully observed. The data audit checks trace order, duplicated event IDs and cross-application offer references; none of those conflicts were found.

van Dongen, B. (2017): BPI Challenge 2017. Eindhoven University of Technology / 4TU.ResearchData. https://doi.org/10.4121/uuid:5f3067df-f10b-45da-b98b-86ae4c7a310b Publisher record ↗

Aggregate research retains attribution and the dataset’s 4TU General Terms of Use (2016), including noncommercial use. No raw records are redistributed. The publisher does not endorse the analysis.

05 / METHOD & LIMITATIONS

Model remaining survival without borrowing future events.

The baselines estimate survival from all training cases or cases with the same stage and age band. A discrete-time boosted hazard model adds event history and elapsed gaps. It learns the probability of reaching the target status in each of the next 30 days.

Validation selects 15 leaves from a fixed 7-versus-15 comparison. The selected model and a reduced stage/age-only model are then fitted on final training. Category mappings are learned within training, and censored partial days contribute only completed risk intervals. Remaining-time means sum the saved survival probabilities; they do not assume an unobserved tail.

  • One institution and one final arrival month; temporal generalization is not established.
  • Early prefixes may omit later loops; unavailable prefixes and already-dispositioned cases are excluded with counts.
  • All predictions and bands refer to a30-day restricted endpoint, not unrestricted remaining duration.
  • Repeated prefixes are dependent and treated as application clusters.
  • No resource-efficiency, causal automation, financial-saving or lending-decision claim is made.
  • Independent technical review is pending.

The public interaction uses saved predictions and explicitly assumed proportional capacity arithmetic. It is not a queueing model or evidence that automation improves throughput.

Frozen endpoint, split and evaluation protocol ↗

06 / CODE

Trace every finding to a saved run.

Run S58-9dd9fde5-183c5e51
Analysis commit 9dd9fde5837bd032b6f3a13c5ae1e2ba7d768a05

git clone https://github.com/mpgibb/michael-gibb-research.git
cd michael-gibb-research
uv sync --frozen
uv run python -W error studies/S58/study.py
uv run python scripts/report_s58.py
uv run python -W error -m unittest discover -s tests -v

Fixed seeds · locked dependencies · complete source checksum. Checks cover future-event exclusion, censored risk intervals, probability identities, cohort accounting and independent reconstruction of final metrics. Independent technical review is pending.