Skip to main content

Benchmark profile

Tool-Agent-User Benchmark (TAU-bench)

Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.

Data verified

How to read this benchmark

Editorial review by Glevd · 2026-07-15

Read this route as the owner of original 2024 TAU-bench methodology. Compare a result only when the domain, archived task release, agent and user models, prompting strategy, tools, number of trials, and pass^k metric are identified. Do not merge it with τ² or τ³ results.

Operator receipt: 0 sourced rows are currently displayable on this page.

Honest limit: BenchLM stores 38 raw numbers under the original key, but none has an exact source attachment or the domain and pass^k labels needed for a valid comparison. The route therefore publishes no score table. The archived repository also warns that its airline and retail tasks are outdated.

About TAU-bench

Year

2024

Tasks

Airline and retail task sets in the archived 2024 release

Format

Domain-specific pass^1 through pass^4 task success

Difficulty

Policy-constrained, multi-turn customer service

The original release reports airline and retail separately and measures reliability with pass^1 through pass^4 across repeated trials. Its repository now warns that those task files are outdated and directs evaluators to the maintained τ³-bench repository.

BenchLM freshness & provenance

Version

TAU-bench 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

Why is there no original TAU-bench score table?

The 38 stored raw numbers do not have exact source attachments or the domain and pass^k labels required for a valid comparison. BenchLM withholds them until those receipts are attached instead of presenting unsupported values as a leaderboard.

Is original TAU-bench still the current release?

No. The original repository warns that its airline and retail task files are outdated and points evaluators to the maintained τ³-bench repository for corrected tasks and newer evaluation modes.

What does original TAU-bench measure?

It measures whether an agent can complete airline or retail customer-service tasks through multi-turn conversation and domain tools while following policy. The paper reports domain-specific pass^k reliability across repeated trials.

Last updated: July 23, 2026 · BenchLM version TAU-bench 2024

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.