Convex Markets / Datasets / RL environments / tau-bench (τ-bench)
SKU RLE-1206 · Sold by External

tau-bench (τ-bench)

Product specifications

SKURLE-1206
Data typeRL environments
Volume165 tasks (retail 115 + airline 50)
Size on diskNot published by source
FormatPython package + JSON domain data
Access modelPUBLIC LICENSE
PricingFree · open-source license
Quality score
LicenseMIT
τ-bench evaluates language agents in dynamic customer-service conversations across two domains — retail (115 tasks) and airline (50 tasks) — where the agent must follow a domain policy document while calling domain-specific API tools and responding to an LLM-simulated user. Agents perform read/write database actions (e.g., modify_reservation, cancel) and are scored by comparing the final database/output state, reported with pass@k (pass^1..pass^k) reliability metrics. Distributed as a pip-installable Python package (MIT-licensed) by Sierra Research.