S59 / Professional services and enterprise AI
Complaint triage with evidence and a human-review threshold
A customer-operations team needs consistent routing of complaints while recognizing uncertain or emerging issues. Compare a simple text classifier with an evidence-grounded language-model workflow under a fixed human-review budget.
Planned topic
Proposed design & source
Selective classification; hierarchical labels; evidence attribution; drift
Use eligible public CFPB narratives and harmonized product/issue labels. Predict intake categories from narrative text only; exclude recorded outcome and routing fields from inputs. Add structured issue extraction with cited spans and an abstention rule. Treat historical category labels as imperfect supervision and document publication selection.
Evaluation: Hold out later periods and a company-transfer slice. Compare TF-IDF/logistic and language-model approaches using macro-F1, rare-issue recall, calibration, coverage-risk curves, latency and cost. Audit factual support on an independently reviewed sample; separate weak-label accuracy from adjudicated quality.
Boundary: Complaints are selected reports, not verified facts or a customer-base denominator. Do not infer company misconduct or rank complaint rates without exposure; no live customer data or automated external action is needed.
Consumer Complaint Database ↗
S60 / Professional services and enterprise AI
Can an AI service agent complete the task reliably?
An enterprise AI leader must choose an agent design based on reliable task completion, policy adherence and cost. Test whether explicit state tracking, tool validation and escalation improve repeated task success over a basic tool-calling agent.
In progress
Proposed design & source
Sequential decisions; state invariants; selective automation; pass^k
Pin the original tau-bench repository, environment, task split, model versions and user-simulator configuration. Compare a baseline agent with one intervention at a time: structured task state, validated tool arguments, policy checks or escalation. Run repeated trials on identical task sets; isolate all tools in the benchmark's simulated environment.
Evaluation: Reserve tasks for final evaluation and use matched trial seeds where possible. Report goal-state success, pass^k, policy violations, unnecessary tool calls, escalation, latency and cost with task-clustered uncertainty. Validate final database state rather than relying only on an LLM judge's impression.
Boundary: Simulated customer-service tasks do not establish real customer satisfaction or labor savings. Escalation is not autonomous completion; pin benchmark versions and distinguish synthetic stress tasks from the official evaluation.
tau-bench ↗
Research code ↗