S49 / Media and entertainment
Personalization beyond popularity: accuracy, diversity and cold start
A media product team wants useful recommendations without showing everyone the same popular titles. Test how much ranking quality must be traded for catalog coverage, diversity and performance on users with little observed history.
Planned topic
Proposed design & source
Matrix factorization; multiobjective ranking; selection-bias analysis
Create global chronological train/validation/test periods from MovieLens ratings. Compare popularity and neighborhood baselines with matrix factorization and a diversity-aware reranker. Define relevance from observed ratings and document candidate construction; use short-history cohorts based only on information available at the cutoff.
Evaluation: Report NDCG/Recall under a fixed documented protocol, rating error where appropriate, catalog coverage, novelty and user-history subgroup results. Compare full eligible-catalog and sampled-candidate sensitivity where computationally feasible; keep future interactions out of popularity statistics.
Boundary: Offline ratings do not demonstrate watch time, subscription retention or causal engagement lift. Respect MovieLens's research-use terms; the publisher requires permission for commercial use, so public demo contents must be checked before release.
MovieLens 32M ↗
S50 / Media and entertainment
News recommendation when freshness and relevance compete
A news product team needs to surface relevant articles as interests and the news cycle change. Test whether content-aware user-history models improve ranking, especially for newly encountered articles, without excessively narrowing topic exposure.
Planned topic
Proposed design & source
Semantic encoders; sequential preferences; diversity-aware reranking
Use MIND's impression candidates, observed clicks and preceding user histories. Compare popularity/content-similarity baselines with a news-encoder and user-history model. Preserve supplied temporal boundaries and evaluate only candidates available in each impression; begin with MIND-small before scaling.
Evaluation: Measure impression-level NDCG, MRR and calibration plus cold-article performance and topic coverage. Compare history lengths and diversity penalties, using user-clustered uncertainty where appropriate. Do not equate a non-click on an exposed candidate with dislike of every unexposed article.
Boundary: The benchmark cannot establish live CTR lift, reader welfare or subscriber revenue. Use only licensed text fields and approved display content; do not imply unrestricted redistribution of full news articles.
MIND — Microsoft News Dataset ↗
S51 / Media and entertainment
Customer-review intelligence that anticipates emerging product issues
A media merchandising or product team needs to detect recurring customer concerns before they dominate a category. Test whether prior-period review themes predict a later increase in negative-review share better than star-rating trends alone.
Planned topic
Proposed design & source
Topic models; structured NLP; hierarchical rates; temporal validation
Choose a reproducible Books, Movies/TV or Digital Music subset. Extract a restrained theme taxonomy using interpretable topic models and a structured language-model challenger. Build item/category-period summaries from reviews already available; reserve later review periods for evaluation and group related products where identifiable.
Evaluation: Compare rating-only and theme-augmented temporal forecasts. Report calibration/error for later negative-review share, topic stability and extraction precision on an independently adjudicated sample. Without adjudication, label extraction quality as unverified and publish only benchmark-label performance.
Boundary: Reviews do not reveal complete sales, exposure or the causal effect of a product change. Avoid full-text redistribution and never treat an LLM's fluent theme explanation as independent evidence.
Amazon Reviews 2023 — media subsets ↗