← Home

RESEARCH BY INDUSTRY

Evidence for the decision.

Find published research and synthetic demonstrations, with planned and developing topics clearly separated.

9 published public-data studies4 synthetic demonstrations51 planned or developing topics

Professional services and enterprise AI

1 published study · 2 planned or developing topics

Published findings & demonstrations

Planned topics & work in progress

These topics are separate from completed findings. Each status reflects the work actually performed.

S59 / Professional services and enterprise AI

Complaint triage with evidence and a human-review threshold

A customer-operations team needs consistent routing of complaints while recognizing uncertain or emerging issues. Compare a simple text classifier with an evidence-grounded language-model workflow under a fixed human-review budget.

Planned topic

Proposed design & source

Selective classification; hierarchical labels; evidence attribution; drift

Use eligible public CFPB narratives and harmonized product/issue labels. Predict intake categories from narrative text only; exclude recorded outcome and routing fields from inputs. Add structured issue extraction with cited spans and an abstention rule. Treat historical category labels as imperfect supervision and document publication selection.

Evaluation: Hold out later periods and a company-transfer slice. Compare TF-IDF/logistic and language-model approaches using macro-F1, rare-issue recall, calibration, coverage-risk curves, latency and cost. Audit factual support on an independently reviewed sample; separate weak-label accuracy from adjudicated quality.

Boundary: Complaints are selected reports, not verified facts or a customer-base denominator. Do not infer company misconduct or rank complaint rates without exposure; no live customer data or automated external action is needed.

Consumer Complaint Database ↗

S60 / Professional services and enterprise AI

Can an AI service agent complete the task reliably?

An enterprise AI leader must choose an agent design based on reliable task completion, policy adherence and cost. Test whether explicit state tracking, tool validation and escalation improve repeated task success over a basic tool-calling agent.

In progress

Proposed design & source

Sequential decisions; state invariants; selective automation; pass^k

Pin the original tau-bench repository, environment, task split, model versions and user-simulator configuration. Compare a baseline agent with one intervention at a time: structured task state, validated tool arguments, policy checks or escalation. Run repeated trials on identical task sets; isolate all tools in the benchmark's simulated environment.

Evaluation: Reserve tasks for final evaluation and use matched trial seeds where possible. Report goal-state success, pass^k, policy violations, unnecessary tool calls, escalation, latency and cost with task-clustered uncertainty. Validate final database state rather than relying only on an LLM judge's impression.

Boundary: Simulated customer-service tasks do not establish real customer satisfaction or labor savings. Escalation is not autonomous completion; pin benchmark versions and distinguish synthetic stress tasks from the official evaluation.

tau-bench ↗

Research code ↗

The 60-topic program spans 20 industries, with 9 studies published. Synthetic demonstrations are counted separately. Unpublished topics have no public case-study route.