← Home

RESEARCH BY INDUSTRY

Evidence for the decision.

Find published research and synthetic demonstrations, with planned and developing topics clearly separated.

9 published public-data studies4 synthetic demonstrations51 planned or developing topics

Software and SaaS

0 published studies · 3 planned or developing topics

Planned topics & work in progress

No findings are published for this selection yet. These topics are separate from completed findings. Each status reflects the work actually performed.

S07 / Software and SaaS

Subscription renewal risk with actionable lead time

A subscription business needs enough advance warning to act on likely non-renewals. Test how much predictive value comes from usage deterioration versus payment/renewal history at realistic intervention cutoffs.

Awaiting prerequisite

Proposed design & source

Temporal landmarking; calibrated boosting; discrete-time hazards

Build member snapshots before expected subscription expiry and reproduce the official churn definition. Compare an elastic-net logistic baseline with gradient boosting and, where complete renewal episodes permit, a discrete-time hazard model. Exclude transactions or usage recorded after each prediction cutoff.

Evaluation: Use month-based holdouts with label-maturation gaps. Report log loss, calibration, precision/recall at fixed contact capacity and performance by tenure and plan. Compare several warning horizons and remove payment/usage feature families in ablations.

Boundary: KKBox is a consumer music service, not a B2B account-revenue dataset. Saved subscribers and retention ROI require an intervention experiment; observed churn prediction cannot establish them.

KKBox Churn Prediction Challenge ↗

S08 / Software and SaaS

Does relational learning justify its complexity?

A product analytics leader must choose between engineered SQL features and relational machine learning. Ask whether learning across users, posts and interactions improves future engagement prediction enough to justify maintenance and serving complexity.

Planned topic

Proposed design & source

Relational inductive bias; temporal heterogeneous graphs; boosting

Select a documented rel-stack engagement task and preserve its official prediction horizon and temporal splits. Compare a transparent SQL aggregation plus boosted-tree baseline with a heterogeneous graph model. Restrict every neighbor, edge and aggregate to information available at the task reference time.

Evaluation: Report the official benchmark metric plus probability calibration, resource usage and performance on users with sparse history. Ablate relation types and compare against identical-budget SQL baselines. Audit neighborhood timestamps to expose graph leakage.

Boundary: Q&A engagement is a product-usage proxy. It does not establish paid SaaS revenue, conversion or the causal effect of a customer-success action.

RelBench rel-stack ↗

S09 / Software and SaaS

Developer ecosystem health beyond star counts

A developer-platform or open-source program leader needs to distinguish brief attention from durable participation. Predict repeat contribution and contributor concentration from early public activity patterns.

Planned topic

Proposed design & source

Survival analysis; cohort effects; hierarchical shrinkage

Construct repository and contributor cohorts from documented public event types. Define a qualifying contribution and a future return window in advance; exclude obvious automation in a sensitivity analysis. Model time to the next qualifying contribution with survival models and repository-level partial pooling.

Evaluation: Train on earlier cohorts and evaluate later cohorts, with a repository-held-out stress test. Compare against stars, recent event count and simple recency. Report return-risk calibration, time-dependent prediction error and sensitivity to bot rules and archive coverage.

Boundary: Public GitHub activity excludes private work and paid usage. It cannot establish software sales, organization-wide productivity or the causal effect of community initiatives.

GH Archive ↗

The 60-topic program spans 20 industries, with 9 studies published. Synthetic demonstrations are counted separately. Unpublished topics have no public case-study route.