Benchmark profile
Tool-Agent-User Benchmark (TAU-bench)
Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.
Data verifiedHow to read this benchmark
Editorial review by Glevd · 2026-07-15
Read this route as the owner of original 2024 TAU-bench methodology. Compare a result only when the domain, archived task release, agent and user models, prompting strategy, tools, number of trials, and pass^k metric are identified. Do not merge it with τ² or τ³ results.
Operator receipt: 0 sourced rows are currently displayable on this page.
Honest limit: BenchLM stores 38 raw numbers under the original key, but none has an exact source attachment or the domain and pass^k labels needed for a valid comparison. The route therefore publishes no score table. The archived repository also warns that its airline and retail tasks are outdated.
About TAU-bench
Year
2024
Tasks
Airline and retail task sets in the archived 2024 release
Format
Domain-specific pass^1 through pass^4 task success
Difficulty
Policy-constrained, multi-turn customer service
The original release reports airline and retail separately and measures reliability with pass^1 through pass^4 across repeated trials. Its repository now warns that those task files are outdated and directs evaluators to the maintained τ³-bench repository.
BenchLM freshness & provenance
Version
TAU-bench 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
Why is there no original TAU-bench score table?
The 38 stored raw numbers do not have exact source attachments or the domain and pass^k labels required for a valid comparison. BenchLM withholds them until those receipts are attached instead of presenting unsupported values as a leaderboard.
Is original TAU-bench still the current release?
No. The original repository warns that its airline and retail task files are outdated and points evaluators to the maintained τ³-bench repository for corrected tasks and newer evaluation modes.
What does original TAU-bench measure?
It measures whether an agent can complete airline or retail customer-service tasks through multi-turn conversation and domain tools while following policy. The paper reports domain-specific pass^k reliability across repeated trials.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.