← Home

APPLIED RESEARCH

A business question.
A complete body of evidence.

Start with the decision and finding. Explore the methods, uncertainty and reproducible source. Browse the research agenda for questions still being investigated.

Find the research behind the decision.

6 of 60 program studies published

4 evaluated synthetic demonstrations

20 industries in the research agenda

54 matching studies

S01 / Retail and e-commerce

Customer value and promotion concentration

A retail marketing director must decide which customer segments deserve limited campaign capacity. Test whether forward-looking customer value identifies future high-value households better than recent spending alone, while detecting dependence on discounts.

Planned

Hierarchical count/spend models; gradient boosting; customer value

Research question and proposed design

Create household-month snapshots from transactions, baskets and available promotion records. Predict next-quarter purchasing and net sales using only prior history. Model transaction frequency and basket value separately; compare a hierarchical count/spend model with gradient boosting. Track discounted and full-price purchasing as separate outcomes.

Evaluation: Use rolling quarterly holdouts and household-clustered uncertainty. Compare against recency-frequency-monetary scoring and last-quarter sales. Report forecast error, calibration by value decile, top-budget value capture, and sensitivity to returns and discount accounting.

Boundary: The frequent-shopper panel is selective. Promotion exposure is observational: the project cannot establish incremental sales caused by a coupon or prove retention ROI.

No completed finding is claimed for this topic.

dunnhumby — The Complete Journey ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S03 / Retail and e-commerce

When does a customer become worth winning back?

An e-commerce team needs to distinguish temporarily inactive buyers from customers unlikely to return. Estimate future transaction counts and revenue distributions to support the timing and prioritization of win-back campaigns.

Planned

BG/NBD; Gamma-Gamma; count/spend challenger

Research question and proposed design

Clean cancellations, returns and missing identifiers under documented rules. Fit BG/NBD purchase-frequency models and a Gamma-Gamma spend model on eligible positive purchases; compare with a flexible count/spend model when assumptions fail. Use successive purchase-history cutoffs and future observation windows.

Evaluation: Compare against simple recency rules and RFM segments. Evaluate held-out transaction-count error, aggregate revenue bias, calibration by recency/frequency cohort and interval coverage. Test whether frequency and spend independence is plausible; publish the challenger if it performs better.

Boundary: The data contain no randomized win-back intervention. Long-term value extrapolation is assumption-sensitive and concerns an older single retailer; do not present revenue as profit.

No completed finding is claimed for this topic.

Online Retail II ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S05 / Advertising and marketing technology

How fragile is multi-touch attribution?

A growth team wants to understand how campaign-credit rules change with attribution windows and incomplete journeys. Study whether sequence-aware conversion forecasts improve prediction and whether channel/touch credit is stable enough to inform further experiments.

Planned

Discrete-time hazards; sequence learning; attribution sensitivity

Research question and proposed design

Reconstruct eligible user journeys from timestamped impressions and clicks. Compare last-touch and time-decay summaries with a discrete-time conversion-hazard model and a compact sequence model. Apply the same observation and conversion windows across methods; use only available anonymized touch identifiers.

Evaluation: Use chronological cutoffs, a gap for delayed conversions and user-grouped sensitivity checks. Report log loss, calibration, time-to-conversion error where identifiable, and credit-rank stability under window changes, touch deletion and identity fragmentation. Compare to a no-history baseline.

Boundary: Thirty days of observational, anonymized advertising records do not reveal causal channel lift. Do not convert attribution credit into verified incremental ROAS or invent absent spend fields.

No completed finding is claimed for this topic.

Criteo Attribution Modeling for Bidding ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S06 / Advertising and marketing technology

The accuracy–latency frontier in ad prediction

An advertising platform must score high-volume traffic within a serving budget. Determine which model delivers the best calibrated click predictions per unit of memory, latency and training cost.

Planned

Sparse logistic regression; factorization machines; Pareto analysis

Research question and proposed design

Use a documented chronological subset first, then a larger scale tier. Compare hashed logistic regression, a factorization machine and one nonlinear interaction model. Keep feature transformations identical where possible; handle unseen categories and changing feature frequencies explicitly.

Evaluation: Hold out later daily files. Measure log loss, precision-recall performance, calibration, peak memory, throughput and p50/p95 latency under a fixed hardware and batch-size protocol. Compare sample-size scaling and ablate interaction features; include cold-start and drift slices.

Boundary: Click labels do not measure incremental sales or advertising profit. Report measured benchmark compute costs separately from extrapolated production costs; do not download or train on the full terabyte by default.

No completed finding is claimed for this topic.

Criteo 1TB Click Logs ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S07 / Software and SaaS

Subscription renewal risk with actionable lead time

A subscription business needs enough advance warning to act on likely non-renewals. Test how much predictive value comes from usage deterioration versus payment/renewal history at realistic intervention cutoffs.

Awaiting prerequisite

Temporal landmarking; calibrated boosting; discrete-time hazards

Research question and proposed design

Build member snapshots before expected subscription expiry and reproduce the official churn definition. Compare an elastic-net logistic baseline with gradient boosting and, where complete renewal episodes permit, a discrete-time hazard model. Exclude transactions or usage recorded after each prediction cutoff.

Evaluation: Use month-based holdouts with label-maturation gaps. Report log loss, calibration, precision/recall at fixed contact capacity and performance by tenure and plan. Compare several warning horizons and remove payment/usage feature families in ablations.

Boundary: KKBox is a consumer music service, not a B2B account-revenue dataset. Saved subscribers and retention ROI require an intervention experiment; observed churn prediction cannot establish them.

No completed finding is claimed for this topic.

KKBox Churn Prediction Challenge ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S08 / Software and SaaS

Does relational learning justify its complexity?

A product analytics leader must choose between engineered SQL features and relational machine learning. Ask whether learning across users, posts and interactions improves future engagement prediction enough to justify maintenance and serving complexity.

Planned

Relational inductive bias; temporal heterogeneous graphs; boosting

Research question and proposed design

Select a documented rel-stack engagement task and preserve its official prediction horizon and temporal splits. Compare a transparent SQL aggregation plus boosted-tree baseline with a heterogeneous graph model. Restrict every neighbor, edge and aggregate to information available at the task reference time.

Evaluation: Report the official benchmark metric plus probability calibration, resource usage and performance on users with sparse history. Ablate relation types and compare against identical-budget SQL baselines. Audit neighborhood timestamps to expose graph leakage.

Boundary: Q&A engagement is a product-usage proxy. It does not establish paid SaaS revenue, conversion or the causal effect of a customer-success action.

No completed finding is claimed for this topic.

RelBench rel-stack ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S09 / Software and SaaS

Developer ecosystem health beyond star counts

A developer-platform or open-source program leader needs to distinguish brief attention from durable participation. Predict repeat contribution and contributor concentration from early public activity patterns.

Planned

Survival analysis; cohort effects; hierarchical shrinkage

Research question and proposed design

Construct repository and contributor cohorts from documented public event types. Define a qualifying contribution and a future return window in advance; exclude obvious automation in a sensitivity analysis. Model time to the next qualifying contribution with survival models and repository-level partial pooling.

Evaluation: Train on earlier cohorts and evaluate later cohorts, with a repository-held-out stress test. Compare against stars, recent event count and simple recency. Report return-risk calibration, time-dependent prediction error and sensitivity to bot rules and archive coverage.

Boundary: Public GitHub activity excludes private work and paid usage. It cannot establish software sales, organization-wide productivity or the causal effect of community initiatives.

No completed finding is claimed for this topic.

GH Archive ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S10 / Logistics and supply chain

Delivery sequencing that combines optimization and driver experience

A last-mile operator needs feasible routes that balance travel efficiency with practical delivery order. Test whether learning from historical high-quality routes improves a constraint-based routing baseline.

Planned

Constrained routing; learning-to-rank; inverse-optimization ideas

Research question and proposed design

Use the provided travel-time matrix, stop/package features and documented constraints. Compare nearest-neighbor and classical local-search baselines with learned sequence preferences inside a constrained routing solver. Preserve an untouched route-level benchmark split and use station holdouts as a transfer test where feasible.

Evaluation: Use the official sequence-quality metric, constraint-violation counts, supplied travel-time totals and computation time. Ablate learned preferences and test travel-time perturbations. Show when a shorter route deviates from expert order rather than treating historical behavior as a globally optimal solution.

Boundary: Locations are obfuscated and route labels do not measure realized wage or fuel savings. Show geographic displays as schematic when appropriate and keep cost conversions explicitly assumed.

No completed finding is claimed for this topic.

Amazon Last Mile Routing Research Challenge ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S11 / Logistics and supply chain

Delivery promises and early exception prioritization

An e-commerce operations team needs to identify orders likely to miss their delivery promise early enough to respond. Estimate both lateness probability and a delivery-time distribution at order approval.

Planned

Quantile boosting; survival/competing events; capacity-aware triage

Research question and proposed design

Join orders, items, seller information and coarse origin/destination geography with explicit grain checks. Use only fields available at approval; compare quantile boosting with a censoring-aware time-to-delivery model. Separate cancellation from delivery rather than treating undelivered orders as on-time.

Evaluation: Use chronological order holdouts and seller-held-out checks. Compare against the promised date and historical lane medians. Report interval coverage, late-order precision/recall at queue capacity and regional calibration; treat later reviews as outcomes, never predictors.

Boundary: The data do not show the causal effect of proactive contact or expedited shipping. Respect the dataset's noncommercial/share-alike terms and avoid publishing raw customer records.

No completed finding is claimed for this topic.

Brazilian E-Commerce Public Dataset by Olist ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S12 / Logistics and supply chain

Regional freight resilience under capacity shocks

A supply-chain planning leader must understand dependence on particular corridors, modes and regions. Identify where a disruption concentrates exposure and compare hypothetical diversification plans.

Planned

Network flows; robust optimization; CVaR stress analysis

Research question and proposed design

Build a commodity-specific origin-destination flow network from a pinned FAF release. Distinguish estimated base-year flows from projections. Formulate an interregional transportation problem with explicitly assumed modal capacities, transfer penalties and demand scenarios; add physical infrastructure data only as a separately documented extension.

Evaluation: Compare current shares, proportional diversion and optimized diversification under the same scenarios. Verify conservation and feasibility; report unserved tonnage, concentration and scenario cost. Test rank stability across releases and commodity definitions; validate projections against later observations only where comparable vintages exist.

Boundary: FAF is an estimated aggregate flow system, not shipment traces or a physical road network. Capacity, rerouting feasibility and costs cannot be inferred from tonnage alone.

No completed finding is claimed for this topic.

Freight Analysis Framework ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S14 / Manufacturing

Maintenance warnings without an alarm flood

A maintenance team needs warnings early enough to schedule work, with few false alarms. Compare interpretable vibration change detection against an unsupervised anomaly model across independent bearing experiments.

Planned

Spectral features; control charts; change-point detection

Research question and proposed design

Extract spectral-band energy, envelope features and robust distribution summaries from vibration windows. Define an early-run reference period explicitly rather than assuming verified healthy labels. Compare robust control charts/change-point methods with a compact anomaly detector and enforce an alarm-persistence rule.

Evaluation: Keep complete experiments or bearings together and show leave-run-out results wherever the small number of runs allows. Report time before documented failure, false alarms per operating hour and sensitivity to window length. Bootstrap by run only when meaningful; emphasize per-run results over spurious narrow intervals.

Boundary: A few laboratory failures cannot establish fleet-wide reliability. The onset of degradation is not fully labeled; maintenance savings and the effect of a repair policy require explicit simulation or prospective data.

No completed finding is claimed for this topic.

IMS Bearings ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S15 / Manufacturing

Remaining useful life with decisions that acknowledge uncertainty

A fleet planner must choose when to replace an asset when failure timing is uncertain. Test whether calibrated remaining-life distributions improve a simulated replacement policy relative to point estimates.

Planned

Degradation modeling; quantile sequences; asymmetric decision loss

Research question and proposed design

Use engine-level C-MAPSS train/test partitions and preserve operating-condition subsets. Compare an engineered degradation-index baseline, boosted quantile regression and one sequence model. Make remaining-life label capping and sensor normalization explicit, including ablations.

Evaluation: Report engine-level RUL error, the benchmark's asymmetric score and interval coverage by operating condition and life stage. Test an unseen-condition subset. Simulate replacement using only sequentially available measurements and compare fixed-age, point-estimate and uncertainty-aware rules.

Boundary: C-MAPSS is simulator-generated. Both asset behavior and the policy economics must be labeled as simulation; no real fleet savings or safety validation are established.

No completed finding is claimed for this topic.

C-MAPSS turbofan degradation ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S16 / Healthcare delivery

Planning for concentrated healthcare expenditure

A healthcare planning team needs to anticipate next-year utilization and spending concentration. Determine whether baseline utilization plus access indicators predicts future expenditure better than demographics alone, without treating high spending as synonymous with clinical need.

Planned

Two-part models; complex survey inference; longitudinal prediction

Research question and proposed design

Use eligible two-year MEPS panel records and the matching longitudinal weights. Predict second-year total expenditure and utilization from first-year information. Compare a two-part expenditure model with a flexible challenger; preserve survey strata and clusters in inference and analyze loss to follow-up.

Evaluation: Hold out later panels where comparable definitions exist. Compare weighted mean, demographic and prior-spending baselines. Report weighted error, top-risk-group calibration, expenditure concentration captured and subgroup uncertainty; bootstrap or linearize according to the survey design.

Boundary: These are population planning estimates, not individual care recommendations. MEPS is not hospital operational telemetry; predicted cost must not be presented as a measure of treatment value or unmet clinical need.

No completed finding is claimed for this topic.

Medical Expenditure Panel Survey (MEPS) ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S17 / Healthcare delivery

Emergency-department waiting-time inequality and uncertainty

A healthcare operations leader wants to know which visit groups experience the longest waits and how stable those patterns are. Estimate adjusted waiting-time distributions, not just a systemwide average.

Planned

Survey-weighted quantiles; standardization; tail-risk analysis

Research question and proposed design

Select NHAMCS emergency-department years with comparable waiting-time, arrival and triage fields after codebook review. Use survey-weighted distributional or quantile models and a long-wait indicator model. Audit missing, capped and special-code values and distinguish visits that were never seen where identifiable.

Evaluation: Compare weighted empirical quantiles with adjusted estimates and leave-year-out predictions. Report uncertainty, effective sample sizes, weighted calibration and sensitivity to missing waits. Suppress unreliable fine-grained cells instead of producing precise-looking rankings.

Boundary: NHAMCS is a sampled visit survey ending in 2022, not a complete hospital arrival/service log. It cannot directly calibrate a particular hospital's queue or prove staffing changes reduce waits.

No completed finding is claimed for this topic.

NHAMCS emergency department public-use files ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S18 / Healthcare delivery

Medicare service concentration and regional market coverage

A healthcare strategy team needs to understand service mix and dependence on a small number of providers. Identify regional concentrations and changing specialty/service patterns that merit a closer market assessment.

Planned

Concentration analysis; matrix factorization; multilevel trends

Research question and proposed design

Construct provider-service-year panels from the public CMS aggregates. Harmonize procedure codes and geographic definitions; distinguish provider location from beneficiary residence. Use service-mix factorization or clustering plus multilevel trend models. Add compatible Medicare enrollment denominators only as an explicitly sourced extension.

Evaluation: Test cluster stability across years and alternative volume definitions. Compare trend forecasts against last-year values, using later years as holdouts. Report suppression-related missingness and sensitivity to geographic boundaries and provider identifiers.

Boundary: Original Medicare Part B is only part of the market. Provider location does not prove patient access, service volume is not clinical quality, and concentration alone does not establish market power.

No completed finding is claimed for this topic.

Medicare Physician & Other Practitioners ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S19 / Pharmaceuticals and biotechnology

Compound selection when laboratory tests are expensive

A discovery team must choose which compounds to test next under a limited assay budget. Evaluate whether uncertainty-aware selection finds active compounds more efficiently than random screening or choosing the largest predicted activity.

Planned

Molecular representations; uncertainty estimation; active learning

Research question and proposed design

Choose a well-populated ChEMBL target and a compatible endpoint. Harmonize units, assay types and duplicate compound measurements before defining the label. Compare molecular-fingerprint models with one graph-based challenger. Run retrospective active-learning rounds by revealing withheld assay labels only after a selection.

Evaluation: Use scaffold-separated and, when feasible, publication-time holdouts. Compare random, greedy and uncertainty/diversity acquisition under identical label budgets. Report hit rate, enrichment, calibration and chemical diversity; repeat acquisition trials with multiple seeds.

Boundary: This is retrospective selection within a measured compound pool. It does not establish clinical efficacy, safety, prospective laboratory hit rates or discovery of a new drug.

No completed finding is claimed for this topic.

ChEMBL ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S20 / Pharmaceuticals and biotechnology

Trial completion risk from information available at registration

A clinical-development operations team must anticipate prolonged or unsuccessful trial execution. Study whether registration-time design and enrollment characteristics predict completion timing and early termination.

Planned

Competing risks; survival calibration; historical snapshot design

Research question and proposed design

Define a cohort by registration period and study type. Reconstruct fields as recorded at registration using accessible record history or dated snapshots; do not use later actual enrollment or revised completion dates as baseline predictors. Fit competing-event models for completion and termination, retaining ongoing trials as censored.

Evaluation: Use later registration cohorts as holdouts, allowing adequate follow-up. Compare simple phase/design strata with regularized survival models. Report horizon-specific calibration, censoring-adjusted Brier scores, event counts and sensitivity to status-reporting lag.

Boundary: Current registry records alone cannot support a leakage-free registration-time forecast. If history cannot be recovered, deliver a retrospective association study and label it accordingly; registry completion is not scientific success.

No completed finding is claimed for this topic.

ClinicalTrials.gov study records ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S21 / Pharmaceuticals and biotechnology

Does molecular data add stable prognostic information?

A translational research team must decide whether a complex molecular signature adds information beyond a simpler clinical baseline. Test this increment under rigorous leakage and batch-effect controls in one cancer cohort.

Planned

Penalized Cox models; pathway aggregation; multi-omic integration

Research question and proposed design

Select only open TCGA clinical and processed molecular files with usable outcome definitions. Start with a clinical Cox model, add a prespecified pathway-level representation and compare with penalized multi-omic survival modeling. Perform normalization and feature selection inside training folds and keep all samples from a patient together.

Evaluation: Use nested patient-level validation and a collection-site holdout if feasible. Report time-specific calibration, concordance, integrated Brier score and feature stability. Test proportional-hazards assumptions and clinical-variable missingness; identify external validation as a future requirement.

Boundary: Retrospective prognostic association does not identify treatment benefit. Open TCGA cohorts do not establish clinical utility or readiness for patient care; controlled-access files are outside this project's scope.

No completed finding is claimed for this topic.

TCGA open-access data through GDC ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S22 / Sports analytics and baseball

Pitcher development signals that survive the next season

A player-development group needs to distinguish durable pitch quality from small-sample outcomes. Test whether pitch characteristics and context improve forecasts of future swing-and-miss and contact quality beyond recent results.

Awaiting prerequisite

Hierarchical shrinkage; staged probabilities; temporal transfer

Research question and proposed design

Create pitcher/pitch-type season histories from Statcast with explicit tracking-era and field-availability checks. Model swing, miss and contact outcomes in stages, using count, handedness and pitch characteristics available at the relevant stage. Pool sparse pitcher/pitch-type estimates hierarchically and forecast a later season.

Evaluation: Use season-forward holdouts and pitcher-level uncertainty. Compare with prior-season rates and league/pitch-type averages. Report calibration, log loss for binary outcomes, contact-quality error and rank stability across pitch-count thresholds and tracking eras.

Boundary: Observed pitch selection depends on opponent and situation. The study cannot prove that changing pitch mix improves performance, and outcome-stage variables must not leak into pre-pitch forecasts.

No completed finding is claimed for this topic.

Statcast / Baseball Savant ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S23 / Sports analytics and baseball

Baseball strategy as a risk-sensitive decision problem

A baseball strategy analyst wants to quantify when a stolen-base attempt or other base-running choice is worth the risk. Estimate the break-even success probability under different base-out states and scoring environments.

Planned

Markov reward processes; expected utility; partial pooling

Research question and proposed design

Parse Retrosheet events into validated base-out transitions and inning runs. Estimate era-specific run-expectancy tables with shrinkage for sparse states. For a stolen-base decision, compare expected runs after success, failure and no attempt using documented transition assumptions.

Evaluation: Hold out seasons and check run-expectancy calibration and transition validity. Compare pooled and era-specific tables, bootstrap by game, and test rare-state sensitivity. Validate the transition engine with known game-state identities and explicitly distinguish predictive checks from causal policy evaluation.

Boundary: Managers select attempts non-randomly, and observed non-attempts are not randomized controls. Historical decision replay cannot by itself establish that a new strategy would cause more wins.

No completed finding is claimed for this topic.

Retrosheet play-by-play ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S24 / Sports analytics and baseball

Roster allocation under aging and performance uncertainty

A roster planner must allocate a fixed budget across players with uncertain future production. Test whether modeling aging, playing time and downside risk changes a roster relative to ranking last season's statistics.

Planned

Hierarchical aging curves; survivor-bias analysis; portfolio optimization

Research question and proposed design

Build historical player-season panels with era normalization and a prespecified offensive run-value formula. Forecast playing time and offensive production jointly using hierarchical age curves and flexible challengers. Optimize a simplified roster under position and budget constraints, using documented salary-covered seasons or clearly synthetic costs.

Evaluation: Use season-forward holdouts and compare against last-season production and age-neutral forecasts. Report production error, interval coverage and roster-performance distributions in historical replay. Stress retirement/selection assumptions and salary coverage; keep hindsight-optimal comparisons confined to the simulation.

Boundary: Lahman is not a complete current payroll or scouting database. A simplified offensive roster omits defense, injuries and contractual constraints unless explicitly added; it cannot establish current front-office value.

No completed finding is claimed for this topic.

Lahman Baseball Database ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S25 / Fitness and wellness

Daily activity patterns and the limits of wellness segmentation

A wellness analytics team wants to understand whether daily activity rhythms reveal useful population segments beyond total movement. Study how pattern estimates change with wear-time quality and demographic composition.

Planned

Functional PCA; clustering stability; complex survey inference

Research question and proposed design

Build participant-level daily curves from NHANES monitor summaries using documented MIMS units, wear/sleep flags and quality checks. Fit functional principal components and a small, stable clustering model. Link same-cycle public demographics and selected examination/laboratory measures by participant ID only as explicitly documented companion files.

Evaluation: Compare rhythm-based segments against total-activity bins. Assess within-person day-to-day reliability, weighted cluster stability and sensitivity to valid-day thresholds. Hold participants together and test transfer between survey cycles; report weighted uncertainty and sample exclusions.

Boundary: Use thresholds appropriate to MIMS, not legacy accelerometer-count cutoffs. These data do not measure gym retention or validate personalized wellness advice; linked health endpoints require compatible survey weights and eligibility.

No completed finding is claimed for this topic.

NHANES Physical Activity Monitor ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S26 / Fitness and wellness

Wearable activity recognition that generalizes to a new person

A wearable-product team must balance activity-recognition quality with the burden of additional sensors. Determine which sensor combination maintains useful recognition on people the model has never seen.

Planned

Temporal CNNs; domain generalization; sensor ablation

Research question and proposed design

Segment PAMAP2 streams with documented overlap and label-transition rules. Compare handcrafted features plus a tree model with a compact temporal convolutional network. Run sensor-location and heart-rate ablations, keeping every window from a participant inside the same outer validation fold.

Evaluation: Use nested leave-one-participant-out validation. Report macro-F1, class-specific confusion, transition-period errors, calibration and per-participant results. Compare sensor subsets under the same protocol and measure latency/model size on stated hardware.

Boundary: Nine participants provide limited evidence about real-world users. Measured inference cost does not directly establish battery life, and the model is an activity benchmark rather than a validated medical monitor.

No completed finding is claimed for this topic.

PAMAP2 Physical Activity Monitoring ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S27 / Fitness and wellness

When more wellness sensors do not mean more evidence

A digital-wellness research team must decide whether combining sleep, activity and cardiac features is promising enough to justify a larger study. Test a deliberately small model of next-morning self-reported affect.

Planned

Bayesian shrinkage; measurement error; incremental predictive value

Research question and proposed design

Use MMASH's short monitoring period and prespecify only a few interpretable features observed before the morning questionnaire. Compare prior affect and a mean-only baseline with strongly regularized models adding sleep/activity or heart-rate-variability summaries. Keep the outcome timing and signal-quality exclusions explicit.

Evaluation: Use participant-held-out prediction with training-fold preprocessing; report absolute error, interval width and sensitivity to individual participants and priors. Compare modality additions one at a time. Publish an inconclusive or negative result if extra sensors do not improve held-out prediction.

Boundary: The study contains 22 healthy young adult men monitored for roughly one day. It cannot support longitudinal recovery forecasting, broad consumer claims or clinical diagnosis; wide uncertainty is a central finding to communicate.

No completed finding is claimed for this topic.

MMASH ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S29 / Banking and fintech

Mortgage application outcomes across markets and lenders

A lending strategy and model-risk team wants to understand geographic variation in application outcomes and which differences remain after measured application characteristics are considered.

Planned

Multilevel models; standardization; selection-bias sensitivity

Research question and proposed design

Create comparable HMDA cohorts by year, loan purpose, product and application disposition. Keep withdrawals and incomplete files distinct from denials. Use multilevel logistic models and standardized outcome comparisons; audit privacy-modified, censored and missing fields before constructing covariates.

Evaluation: Hold out later years and selected lenders/regions. Compare raw versus standardized differences with uncertainty, calibration and sensitivity to covariate sets. Avoid ranking tiny cells; evaluate missingness patterns and changes in reporting definitions.

Boundary: HMDA is not a loan-performance panel and does not contain every underwriting factor. This is an audit and research demonstration, not a legal finding of discrimination or a deployable credit-decision engine.

No completed finding is claimed for this topic.

HMDA mortgage application data ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S30 / Banking and fintech

Credit-risk decisions when calibration matters more than a leaderboard

A risk analytics leader must understand how probability errors affect a hypothetical review policy. Compare interpretable and flexible default models under asymmetric error costs and limited manual-review capacity.

Planned

Proper scoring rules; probability calibration; expected loss

Research question and proposed design

Use the available repayment, bill and payment history to predict the provided next-month default outcome. Fit logistic and monotone/additive baselines plus boosted trees. Reserve an untouched customer-level test set, calibrate on separate validation data and document feature timing.

Evaluation: Use nested development folds and the fixed test set; do not invent a temporal holdout absent multiple cohorts. Report Brier score, log loss, calibration by risk band, PR performance and decision-cost sensitivity. Analyze subgroup errors without making automatic individual lending decisions.

Boundary: This is an older Taiwan cohort with no prospective U.S. validation. Default prediction is not fraud detection; assumed exposure and losses must be separated from observed labels and no actual lending policy is validated.

No completed finding is claimed for this topic.

Default of Credit Card Clients ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S31 / Insurance

Insurance pricing: interpretable structure versus nonlinear accuracy

An insurance analytics team needs accurate expected-loss estimates that remain understandable across policy segments. Test whether a nonlinear model improves frequency-severity predictions while preserving portfolio and segment calibration.

Planned

Exposure-offset GLMs; frequency-severity; Tweedie models

Research question and proposed design

Audit policy exposure and claim counts, aggregate severity by policy before joining, and explain inconsistencies. Compare an exposure-offset Poisson/negative-binomial frequency model plus Gamma severity model with Tweedie and boosted challengers. Keep claim-level observations from the same policy together.

Evaluation: Use policy-held-out development/test sets and geographic stress tests where feasible. Report exposure-weighted deviance, total predicted/observed loss, segment calibration and tail sensitivity. Compare uncapped and transparently capped severity analyses; do not claim a time split without reliable policy dates.

Boundary: Historical French motor data do not validate a current insurance rate filing. Pure premium is expected insured loss, not the final customer price; regulatory and deployment questions are outside the benchmark.

No completed finding is claimed for this topic.

freMTPL2 frequency and severity ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S32 / Insurance

Flood claims and the risk of geographic concentration

An insurer or resilience planner needs to understand tail payments and correlated geographic exposure. Estimate how claim-severity conclusions change when major flood events, coverage limits and portfolio composition are considered.

Planned

Hierarchical severity; extreme-value analysis; event-level dependence

Research question and proposed design

Clean NFIP claim payments and dated event/geography information. Analyze severity conditional on a reported claim using hierarchical models and carefully diagnosed tail models. For incidence or expected loss per insured exposure, explicitly add compatible redacted policy data and define earned-exposure denominators first.

Evaluation: Hold out entire major events or event-years and geographic groups. Compare empirical severity and lognormal/Gamma baselines with tail models. Report tail quantile calibration, threshold sensitivity and event-bootstrap uncertainty; flag unsupported extrapolation beyond observed experience.

Boundary: Claims alone cannot establish flood probability or claim incidence. Redacted geography, evolving limits and nominal payment amounts complicate comparisons; modeled return periods or portfolio losses need explicit additional assumptions.

No completed finding is claimed for this topic.

OpenFEMA NFIP Redacted Claims ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S33 / Insurance

Longevity assumptions and long-horizon liability sensitivity

A benefits or insurance planning team must understand how mortality assumptions affect a long-duration payment obligation. Quantify sensitivity to table choice, longevity improvement and discount rates without pretending the tables are individual training records.

Planned

Life-contingent valuation; survival probabilities; sensitivity analysis

Research question and proposed design

Select compatible SOA tables with clear population, vintage and select/ultimate or generational definitions. Convert death probabilities into survival curves and expected payment streams. Apply a clearly synthetic cohort of ages and benefits, then compare deterministic and explicitly assumed stochastic improvement scenarios.

Evaluation: Check survival monotonicity, probability identities and actuarial present values against hand-calculated cases. Compare suitable table vintages and run rate/improvement stress grids. Do not report predictive accuracy without separate death and exposure observations.

Boundary: This is a transparent scenario engine using published rates and synthetic obligations. It is not a fitted individual mortality model, an audited reserve calculation or a validated forecast of future longevity.

No completed finding is claimed for this topic.

Mortality and Other Rate Tables (MORT) ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S34 / Energy and utilities

Electricity reserve planning from probabilistic load forecasts

A utility planning team must balance excess reserve against costly demand shortfalls. Ask whether calibrated upper-tail demand forecasts produce more reliable reserve decisions than a fixed margin above a point forecast.

Planned

Probabilistic forecasting; chance constraints; asymmetric loss

Research question and proposed design

Forecast hourly balancing-authority demand at a precisely defined issue time using lagged load, calendar information and only legitimately available inputs. Compare seasonal models and quantile boosting; benchmark against published demand forecasts only when their issue time aligns with the task. Weather forecasts are an optional, separately sourced extension.

Evaluation: Use rolling seasonal holdouts and an extreme-demand stress period. Report quantile loss, interval coverage, upper-tail exceedances and modeled reserve/shortfall cost relative to fixed margins. Record data vintages; label revised-data backtests when original releases are unavailable.

Boundary: Demand-only reserve scenarios omit generator outages and full network/security constraints. They are not an operational dispatch plan; do not use realized future weather or mismatched official-forecast timing.

No completed finding is claimed for this topic.

EIA-930 Hourly Electric Grid Monitor ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S35 / Energy and utilities

Demand response: shifting peak load or moving the problem?

A utility wants to know whether time-varying prices reduce peak usage or merely shift it into adjacent periods. Estimate heterogeneous load changes and rebound patterns under the London tariff trial's documented assignment process.

Planned

Intertemporal substitution; panel effects; conditional DiD/ITT

Research question and proposed design

Join household half-hourly consumption with the supplied 2013 price-signal calendar. Audit recruitment, group assignment and pre-period comparability before selecting the estimand. Use randomized assignment analysis only if verified; otherwise use household/time panel models and difference-in-differences with explicit identifying assumptions.

Evaluation: Predefine event windows, peak and adjacent-period outcomes. Compare with flat-tariff households, inspect pre-trends/placebo windows, cluster uncertainty by household and price event as appropriate, and assess heterogeneous effects using held-out households.

Boundary: A tariff-group label alone does not establish random assignment. If design documentation is insufficient, publish adjusted associations and sensitivity analyses rather than causal savings claims.

No completed finding is claimed for this topic.

Low Carbon London SmartMeter data ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S36 / Energy and utilities

Peak-demand prediction across heterogeneous electricity customers

An energy analytics team needs reliable peak forecasts across many customers with different load shapes. Test whether a shared model transfers useful information without sacrificing performance on unusual customers.

Planned

Global/local forecasting; partial pooling; peak quantiles

Research question and proposed design

Build timezone-aware quarter-hourly series with explicit missing-value, zero and daylight-saving rules. Compare seasonal naive forecasts, local statistical models and a global quantile model. Cluster load shapes only within training data and evaluate whether cluster-aware models improve peak forecasting.

Evaluation: Use rolling time holdouts plus customer-held-out transfer tests. Report scaled error, peak-window quantile loss, interval coverage and worst-decile customer performance. Compare aggregate forecasts and sensitivity to DST handling and newly active customers.

Boundary: Customer identities and interventions are largely unknown. This benchmark cannot estimate causal tariff response or infer demographic explanations; capacity lines and operating costs are scenarios unless separately supplied.

No completed finding is claimed for this topic.

ElectricityLoadDiagrams20112014 ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S37 / Airlines and transportation

Early warning of airline disruption across an airport network

An airline operations team must prioritize flights for early disruption review. Test whether recent network conditions add useful warning of severe arrival delays or cancellations beyond schedule and seasonal patterns.

Planned

Temporal network features; calibrated classification; event severity

Research question and proposed design

Define a fixed pre-departure prediction cutoff and create schedule, airport-congestion and lagged completed-flight features. Compare calibrated boosting with a simpler logistic model and a network-aware challenger. Use aircraft rotation features only where the assignment is knowable at the cutoff; exclude the target flight's realized delay causes.

Evaluation: Use rolling month/season holdouts and airport-transfer tests. Report rare-event precision-recall, calibration, warning lead time and severe disruptions captured at a fixed review capacity. Separate cancellation from conditional delay severity and audit publication/schedule revision timing.

Boundary: Retrospective BTS records may not reconstruct every live schedule or aircraft assignment. Do not present hindsight features as real-time knowledge or simulated schedule changes as proven delay reduction.

No completed finding is claimed for this topic.

Airline On-Time Performance ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S38 / Airlines and transportation

Fleet positioning under uncertain urban trip demand

A mobility operator needs to place a limited fleet where upcoming trips are likely. Compare fixed, historically proportional and forecast-driven allocation under explicit travel and repositioning assumptions.

Planned

Spatiotemporal forecasting; min-cost flow; stochastic allocation

Research question and proposed design

Choose one TLC vehicle category and consistent date range. Aggregate completed trips into zone-time demand and estimate travel-time distributions from eligible historical trips. Forecast next-period observed pickups using seasonal and spatially pooled models; feed scenarios into a min-cost-flow allocation model.

Evaluation: Hold out later weeks and demand-shift periods. Report zone-level scaled error, interval coverage and peak-zone performance. In a clearly specified simulator, compare unmet observed-request proxies, relocation distance and modeled cost across policies; vary travel-time and fleet-size assumptions.

Boundary: Completed trips omit unserved requests and empty-vehicle movements. Public monthly files also have publication lag. The project is a historical fleet-planning benchmark, not a validated live dispatch system.

No completed finding is claimed for this topic.

NYC TLC Trip Records ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S39 / Airlines and transportation

Chicago transit planning through structural demand changes

A transit planning team needs a forecast that adapts when commuting patterns change. Test whether regime-aware models improve daily bus and rail boarding forecasts over a stable seasonal baseline.

Planned

State-space models; change detection; coherent aggregation

Research question and proposed design

Use CTA's system-level daily bus, rail and total boarding series. Build calendar/day-type features and compare seasonal naive, dynamic regression and state-space models with change detection. Fit change points using training data only; evaluate known disruption periods as stress tests.

Evaluation: Use rolling 7- and 28-day forecast origins spanning stable and changing demand periods. Report scaled error, interval coverage and recovery after shifts. Compare expanding versus recent-window training and evaluate weekday/weekend performance separately.

Boundary: This specific dataset contains daily system totals, not station-hour demand or individual trips. Capacity scenarios cannot justify train-by-train staffing or prove service changes caused ridership growth without additional data.

No completed finding is claimed for this topic.

CTA Daily Boarding Totals ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S40 / Hospitality and travel

Hotel cancellation risk and the cost of overbooking

A hotel revenue manager must balance empty rooms against the cost of accommodating guests elsewhere. Test whether booking-level cancellation probabilities improve a simulated overbooking policy compared with uniform cancellation assumptions.

Planned

Calibrated cancellation models; revenue management; asymmetric loss

Research question and proposed design

Define a booking-time prediction point and audit whether each field was available then. Exclude final reservation status/date, later booking changes and other post-outcome fields. Compare calibrated logistic/additive models with boosting; aggregate predicted cancellations across arrival dates with dependence sensitivity.

Evaluation: Use time-based arrival holdouts and a hotel-transfer stress test. Report cancellation calibration, log loss and arrival-date prediction error. Compare fixed-rule and model-based overbooking in a simulator with explicit room capacity, room-night logic, rates and displacement costs.

Boundary: Two hotels do not establish universal performance. Original booking-time versions of some fields may be unavailable; label that limitation. Capacity and counterfactual overbooking revenue are simulated, not measured business improvement.

No completed finding is claimed for this topic.

Hotel Booking Demand datasets ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S41 / Hospitality and travel

Short-term rental positioning without invented occupancy

A hospitality market analyst wants to understand which listing features and neighborhoods are associated with asking-price differences. Build a reliable comparison tool that distinguishes a listing's advertised position from actual booking performance.

Planned

Hedonic models; spatial dependence; quantile regression

Research question and proposed design

Choose a city and reproducible snapshot set; normalize currency, minimum-stay rules and listing types. Fit a spatial hedonic model and a quantile-boosting challenger to advertised nightly prices. Group repeated listings across splits and use coarse location features with spatial-block validation.

Evaluation: Compare neighborhood/type medians with model predictions on later snapshots and held-out spatial blocks. Report absolute/log-price error, interval coverage and stability after outlier filtering. Audit calendar availability as a separate recorded field rather than converting it into observed occupancy.

Boundary: Unavailable calendar nights may be booked or blocked. Do not claim measured occupancy, annual rental yield or causal revenue gains from amenities; location and host details should remain appropriately aggregated.

No completed finding is claimed for this topic.

Inside Airbnb ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S42 / Hospitality and travel

Tourism capacity planning across seasons and regions

A tourism or hospitality strategy team needs to distinguish recurring seasonality from lasting changes in demand mix. Forecast accommodation nights and identify regions or visitor segments with the greatest planning uncertainty.

Planned

Seasonal/state-space models; hierarchical forecasting; pooling

Research question and proposed design

Select comparable Eurostat monthly nights/arrivals series at a supported geographic level, separating resident and nonresident demand where available. Harmonize coverage and reporting breaks. Compare seasonal baselines, dynamic regression and globally pooled probabilistic forecasts with consistent totals.

Evaluation: Use rolling forecast origins across normal and disrupted periods. Report scaled error, quantile loss, interval coverage and peak-season bias by region. Compare simple seasonal recovery assumptions with fitted models; include missing-series and definition-change sensitivity.

Boundary: Aggregate nights are not individual bookings or a hotel's revenue. Match the geography and frequency actually available; do not infer property-level occupancy or staffing needs from incompatible regional totals.

No completed finding is claimed for this topic.

Eurostat tourism accommodation statistics ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S44 / Real estate and property management

Housing-market forecasts that acknowledge revisions and turning points

A property strategy team needs to compare regional rent and home-value trajectories without treating estimated indices as executable investment returns. Test whether pooled models detect changing momentum better than simple trend extrapolation.

Planned

Dynamic factors; local trends; real-time/vintage evaluation

Research question and proposed design

Select stable ZHVI/ZORI regional series with documented methodology and available vintages. Compare local trend/seasonal models with a dynamic-factor or pooled forecasting model. Use lagged regional indicators from the same release; introduce outside macro series only with explicit as-of publication alignment.

Evaluation: Use rolling 3-, 6- and 12-month horizons. Report scaled error, directional accuracy, interval coverage and performance around turning points. Compare original-vintage and latest-revised backtests where possible; otherwise label the exercise a revised-data retrospective forecast.

Boundary: Regional indices are modeled aggregates, not individual asset returns, rental cash flows or buy/sell recommendations. Historical revision awareness does not remove uncertainty about future structural changes.

No completed finding is claimed for this topic.

Zillow Research housing data ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S45 / Real estate and property management

Transparent screening of property redevelopment opportunities

A real-estate strategy team needs a defensible way to prioritize parcels for further diligence. Build a transparent screening framework that weighs current use, building characteristics and simple utilization proxies under multiple business preferences.

Planned

Multi-criteria decisions; Pareto dominance; robust ranking

Research question and proposed design

Use PLUTO tax-lot geometry, land use, building age/area and documented zoning-related fields. Validate units and implausible records. Derive a clearly approximate unused-floor-area indicator only where compatible fields permit; use multi-criteria ranking and robust optimization rather than training on an invented redevelopment-profit label.

Evaluation: Compare simple size/age filters with the multi-criteria shortlist. Report rank stability under weight/input perturbations, missing-data sensitivity and independent spot checks against source attributes. If outcome data are later added, define a separate temporal validation study.

Boundary: PLUTO is an inventory, not transaction-return or tenant-demand data. Simple floor-area calculations do not certify buildability; project economics, entitlements and parcel-specific constraints require additional evidence.

No completed finding is claimed for this topic.

PLUTO / MapPLUTO ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S46 / Telecommunications

Balancing retention and expansion in a telecom contact queue

A telecom commercial team must allocate limited contact capacity across churn risk, product appetite and upsell likelihood. Test whether a shared predictive representation improves the three outcomes and clarifies tradeoffs between objectives.

Planned

Multi-task learning; calibrated probabilities; constrained allocation

Research question and proposed design

Use the documented Orange KDD Cup targets and a manageable feature subset before scaling. Compare three independent regularized/boosted models with a shared multi-task representation. Audit missingness and anonymized features; calibrate each target separately and retain the official held-out protocol where labels are available.

Evaluation: Use nested development splits and an untouched customer test; do not invent timestamps. Report target-specific PR metrics, calibration, multi-task gains/losses and capacity-specific outcome capture. Test sensitivity to missing features and assumed value weights.

Boundary: Anonymous benchmark features limit business explanation. These labels do not identify causal offer response, actual revenue or a production-ready next-best-action policy; synthetic value assumptions must be clearly labeled.

No completed finding is claimed for this topic.

Orange KDD Cup 2009 ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S47 / Telecommunications

Service-friction signals and a three-month churn planning window

A telecom service leader wants to understand whether call failures, complaints and usage patterns identify customers at risk before departure. Test stable, interpretable signals and the limits of translating them into an intervention plan.

Planned

Additive models; calibration; causal diagrams; feature stability

Research question and proposed design

Respect the dataset's first-nine-month feature period and end-of-month-12 churn label. Compare regularized logistic/additive models with boosting. Audit status and calculated customer-value fields for circular or late information and report ablations excluding them; avoid treating row order as calendar time.

Evaluation: Use nested stratified development and a fixed customer test. Report PR performance, calibration, recall at service capacity and bootstrap stability of feature effects. Compare a simple complaints/recency-style rule with the learned models and assess dependence on derived fields.

Boundary: A small single-company sample cannot establish that fixing call failures causes retention. There are no repeated time cohorts for a genuine out-of-time validation; inferred customer value is not verified profit.

No completed finding is claimed for this topic.

Iranian Churn ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S48 / Telecommunications

Urban telecom activity and resilient capacity scenarios

A network planning team wants to locate predictable activity peaks and unusual surges. Test whether spatial models improve short-horizon cell-grid activity forecasts and identify where a hypothetical capacity budget would be most exposed.

Planned

Spatiotemporal models; graph regularization; residual anomalies

Research question and proposed design

Build time-aligned Milan grid series for the released activity measures. Compare seasonal baselines, spatially pooled models and a graph-temporal challenger. Define adjacency from the documented grid and preserve separate activity types; treat normalization and missing intervals explicitly.

Evaluation: Use blocked future-day/week holdouts and selected spatial blocks. Report scaled error, peak-period quantile loss and false alerts under a defined threshold. Test the short observation period's sensitivity to unusual days and compare any allocation policies only under common simulated assumptions.

Boundary: Normalized grid activity is not subscriber count, bytes, tower load or dropped-call experience. The short historical period cannot validate a physical network expansion plan or annual seasonality.

No completed finding is claimed for this topic.

Telecom Italia Milan activity ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S49 / Media and entertainment

Personalization beyond popularity: accuracy, diversity and cold start

A media product team wants useful recommendations without showing everyone the same popular titles. Test how much ranking quality must be traded for catalog coverage, diversity and performance on users with little observed history.

Planned

Matrix factorization; multiobjective ranking; selection-bias analysis

Research question and proposed design

Create global chronological train/validation/test periods from MovieLens ratings. Compare popularity and neighborhood baselines with matrix factorization and a diversity-aware reranker. Define relevance from observed ratings and document candidate construction; use short-history cohorts based only on information available at the cutoff.

Evaluation: Report NDCG/Recall under a fixed documented protocol, rating error where appropriate, catalog coverage, novelty and user-history subgroup results. Compare full eligible-catalog and sampled-candidate sensitivity where computationally feasible; keep future interactions out of popularity statistics.

Boundary: Offline ratings do not demonstrate watch time, subscription retention or causal engagement lift. Respect MovieLens's research-use terms; the publisher requires permission for commercial use, so public demo contents must be checked before release.

No completed finding is claimed for this topic.

MovieLens 32M ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S50 / Media and entertainment

News recommendation when freshness and relevance compete

A news product team needs to surface relevant articles as interests and the news cycle change. Test whether content-aware user-history models improve ranking, especially for newly encountered articles, without excessively narrowing topic exposure.

Planned

Semantic encoders; sequential preferences; diversity-aware reranking

Research question and proposed design

Use MIND's impression candidates, observed clicks and preceding user histories. Compare popularity/content-similarity baselines with a news-encoder and user-history model. Preserve supplied temporal boundaries and evaluate only candidates available in each impression; begin with MIND-small before scaling.

Evaluation: Measure impression-level NDCG, MRR and calibration plus cold-article performance and topic coverage. Compare history lengths and diversity penalties, using user-clustered uncertainty where appropriate. Do not equate a non-click on an exposed candidate with dislike of every unexposed article.

Boundary: The benchmark cannot establish live CTR lift, reader welfare or subscriber revenue. Use only licensed text fields and approved display content; do not imply unrestricted redistribution of full news articles.

No completed finding is claimed for this topic.

MIND — Microsoft News Dataset ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S51 / Media and entertainment

Customer-review intelligence that anticipates emerging product issues

A media merchandising or product team needs to detect recurring customer concerns before they dominate a category. Test whether prior-period review themes predict a later increase in negative-review share better than star-rating trends alone.

Planned

Topic models; structured NLP; hierarchical rates; temporal validation

Research question and proposed design

Choose a reproducible Books, Movies/TV or Digital Music subset. Extract a restrained theme taxonomy using interpretable topic models and a structured language-model challenger. Build item/category-period summaries from reviews already available; reserve later review periods for evaluation and group related products where identifiable.

Evaluation: Compare rating-only and theme-augmented temporal forecasts. Report calibration/error for later negative-review share, topic stability and extraction precision on an independently adjudicated sample. Without adjudication, label extraction quality as unverified and publish only benchmark-label performance.

Boundary: Reviews do not reveal complete sales, exposure or the causal effect of a product change. Avoid full-text redistribution and never treat an LLM's fluent theme explanation as independent evidence.

No completed finding is claimed for this topic.

Amazon Reviews 2023 — media subsets ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S52 / Agriculture and food production

Crop-yield uncertainty for procurement planning

A food procurement team needs to anticipate regional supply variability before harvest. Test whether spatially pooled yield forecasts improve uncertainty estimates and hypothetical sourcing allocations over trend-only planning.

Planned

Hierarchical spatial models; trend forecasts; stochastic procurement

Research question and proposed design

Select one major crop and comparable NASS county/state yield series. Define a pre-harvest forecast issue date and use only information published by then. Start with historical trend and lagged yield; treat dated NOAA weather and available acreage estimates as separately documented enhancements, not assumed fields in Quick Stats.

Evaluation: Use year-forward and region-held-out tests. Compare trend, historical-average and pooled models using yield error, interval coverage and adverse-year performance. Audit suppressed/revised estimates. Evaluate procurement regret only in a simulator whose prices, capacity and demand are clearly specified.

Boundary: Public aggregate estimates are not farm-level treatment data. Do not use final acreage, realized future weather or revised releases as if available before harvest, or claim causal effects of agronomic practices.

No completed finding is claimed for this topic.

USDA NASS Quick Stats ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S53 / Agriculture and food production

Crop-rotation forecasting without spatial leakage

An agricultural planning team wants to anticipate regional crop-mix changes and identify where transitions are difficult to forecast. Test whether multi-year rotation history improves next-year crop classification and aggregate acreage estimates.

Planned

Markov transitions; spatial-temporal classification; error propagation

Research question and proposed design

Use past-year Cropland Data Layers to predict the next year's crop class on a fixed, compatible grid. Compare persistence and transition-matrix baselines with a spatial-temporal classifier. Harmonize changing resolution, class definitions and alignment; never use the target-year crop map as an input feature.

Evaluation: Hold out future years and large geographic blocks, not random adjacent pixels. Report class-balanced accuracy, transition-specific confusion and county-level acreage error, with uncertainty clustered spatially. Test sensitivity to map resolution and documented label accuracy.

Boundary: CDL labels are themselves remotely sensed estimates, not perfect field ground truth. This study does not estimate farm profit, yield or the causal benefit of crop rotation; nominal pixel count overstates independent evidence.

No completed finding is claimed for this topic.

USDA Cropland Data Layer ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S54 / Agriculture and food production

Food-supply dependence and trade disruption resilience

A food business or strategy team needs to understand dependence on a small set of supplying countries. Identify concentrated commodity networks and compare conditional sourcing-diversification strategies under production or trade shocks.

Planned

Material-balance networks; concentration; robust optimization

Research question and proposed design

Select compatible FAOSTAT production, trade and food-balance modules for one commodity group. Harmonize physical units, reporting years and product definitions; use bilateral trade tables only where available. Construct supply-dependence indicators and a constrained trade-reallocation scenario model.

Evaluation: Compare concentration-only rankings with network-based exposure. Verify mass-balance compatibility and stress alternative shock magnitudes, substitution limits and data-revision treatments. Compare indicators around historical disruptions descriptively, without claiming the model causally explains observed price changes.

Boundary: Country-level annual flows do not reveal a particular company's contracts, inventory or margin. Reallocation capacity and substitution are scenarios; the tool cannot claim realized savings or precise future food prices.

No completed finding is claimed for this topic.

FAOSTAT ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S55 / Construction and infrastructure

Permit-process duration and project planning risk

A construction planning team needs realistic administrative lead times before committing downstream resources. Estimate the distribution of time from a documented filing milestone to a documented permit milestone and identify where uncertainty is greatest.

Planned

Survival/AFT models; multistate processes; hierarchical pooling

Research question and proposed design

Choose a consistent DOB system and application cohort with reliable milestone dates. Separate application, job and permit identifiers, amendments and withdrawals. Fit hierarchical accelerated-failure-time or survival models using information known at the starting milestone; retain unresolved applications as censored where follow-up is observable.

Evaluation: Hold out later filing cohorts and compare with category-level historical medians. Report horizon-specific calibration, censoring-adjusted error and interval coverage. Audit missing dates, system migrations and whether the apparent endpoint is original issuance or a later revision.

Boundary: Permit dates do not identify construction completion, total project cost or the causal effect of an expediting service. If milestone history is unavailable, narrow the question to the reliably observed administrative interval.

No completed finding is claimed for this topic.

NYC DOB applications and permits ↗

Original dataset selection position 1 within this industry. This is research-selection metadata, not measured model performance.

S56 / Construction and infrastructure

Construction-market outlook for commercial capacity planning

A contractor or construction-services business needs to anticipate shifts in demand across market segments. Test whether pooled sector forecasts improve planning over simple extrapolation of the latest spending growth.

Planned

Dynamic factors; coherent sector forecasts; nominal/real decomposition

Research question and proposed design

Build consistent monthly Census construction-spending series by supported category. Distinguish seasonally adjusted annual rates from unadjusted monthly values and avoid summing incompatible units. Compare local trends, dynamic-factor models and coherent sector forecasts. Add appropriate public price indices only as an explicit real-spending extension.

Evaluation: Use rolling multi-month horizons with seasonal/trend baselines. Report scaled error, turning-point performance, interval coverage and revision sensitivity. Translate forecasts into a hypothetical staffing/capacity scenario only after specifying a business's market share and productivity assumptions.

Boundary: Aggregate spending cannot measure individual project overruns, backlog or contractor profitability. Forecasting a sector does not establish the demand a particular firm will win.

No completed finding is claimed for this topic.

Construction Spending / Value Put in Place ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S57 / Construction and infrastructure

Turning severe-injury narratives into a prevention taxonomy

A construction safety team needs a consistent view of recurring hazard mechanisms across incident reports. Test whether structured text extraction can organize reported incidents accurately enough to support expert review and training priorities.

Planned

Structured NLP; measurement validation; selective classification

Research question and proposed design

Filter the OSHA records to the documented construction industry codes and eligible jurisdiction/time coverage. Build a prespecified mechanism/activity taxonomy. Compare keyword and TF-IDF baselines with a structured language-model classifier; retain supporting text spans and an abstention option for ambiguous cases.

Evaluation: Create a de-identified, independently adjudicated evaluation sample with clear labeling rules; report agreement, class-specific precision/recall and error-versus-review-volume curves. Use later periods for drift checks. Without expert labels, treat the model as an unvalidated organizing aid.

Boundary: Coverage excludes important jurisdictions and the severe-injury report system is not a complete fatality census. Counts lack matching worker-hour denominators; do not rank employers by safety or claim a prevention intervention reduced injuries.

No completed finding is claimed for this topic.

OSHA Severe Injury Reports — construction subset ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

S59 / Professional services and enterprise AI

Complaint triage with evidence and a human-review threshold

A customer-operations team needs consistent routing of complaints while recognizing uncertain or emerging issues. Compare a simple text classifier with an evidence-grounded language-model workflow under a fixed human-review budget.

Planned

Selective classification; hierarchical labels; evidence attribution; drift

Research question and proposed design

Use eligible public CFPB narratives and harmonized product/issue labels. Predict intake categories from narrative text only; exclude recorded outcome and routing fields from inputs. Add structured issue extraction with cited spans and an abstention rule. Treat historical category labels as imperfect supervision and document publication selection.

Evaluation: Hold out later periods and a company-transfer slice. Compare TF-IDF/logistic and language-model approaches using macro-F1, rare-issue recall, calibration, coverage-risk curves, latency and cost. Audit factual support on an independently reviewed sample; separate weak-label accuracy from adjudicated quality.

Boundary: Complaints are selected reports, not verified facts or a customer-base denominator. Do not infer company misconduct or rank complaint rates without exposure; no live customer data or automated external action is needed.

No completed finding is claimed for this topic.

Consumer Complaint Database ↗

Original dataset selection position 2 within this industry. This is research-selection metadata, not measured model performance.

S60 / Professional services and enterprise AI

Can an AI service agent complete the task reliably?

An enterprise AI leader must choose an agent design based on reliable task completion, policy adherence and cost. Test whether explicit state tracking, tool validation and escalation improve repeated task success over a basic tool-calling agent.

Baseline evaluated

Sequential decisions; state invariants; selective automation; pass^k

Research question and proposed design

Pin the original tau-bench repository, environment, task split, model versions and user-simulator configuration. Compare a baseline agent with one intervention at a time: structured task state, validated tool arguments, policy checks or escalation. Run repeated trials on identical task sets; isolate all tools in the benchmark's simulated environment.

Evaluation: Reserve tasks for final evaluation and use matched trial seeds where possible. Report goal-state success, pass^k, policy violations, unnecessary tool calls, escalation, latency and cost with task-clustered uncertainty. Validate final database state rather than relying only on an LLM judge's impression.

Boundary: Simulated customer-service tasks do not establish real customer satisfaction or labor savings. Escalation is not autonomous completion; pin benchmark versions and distinguish synthetic stress tasks from the official evaluation.

No completed finding is claimed for this topic.

tau-bench ↗

Original dataset selection position 3 within this industry. This is research-selection metadata, not measured model performance.

Program progress counts the 60 numbered studies separately from the four original synthetic demonstrations. Planned topics have no public case-study route. Published findings link to evaluated artifacts and research code.