Benchmark profile
τ²-Bench Tool-Agent-User Evaluation (τ²-bench results)
This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.
Data verifiedHow to read this leaderboard
Editorial review by Glevd · 2026-07-15
Use a row only with its attached source and setup label. Match the domain, task release, agent model, user-simulator model, scaffold, prompts, trial count, and pass^k metric before comparing scores. The sorted table is a source ledger, not a controlled cross-provider ranking.
Operator receipt: 145 sourced rows are currently displayable on this page; the highest published score among these 145 models is GLM-5.2 at 99.1%, which does not establish a market leader.
Honest limit: The page mixes a large third-party telecom snapshot with smaller provider-published slices. Those sources do not use one guaranteed-common harness or reporting policy, and BenchLM did not rerun them. A higher number can reflect a different user model, prompt, domain, task release, or repeat policy.
Benchmark score on τ²-bench sourced results — July 23, 2026
BenchLM mirrors the published score view for τ²-bench sourced results. The public snapshot contains 145 models. GLM-5.2 has the highest published score at 99.1%, but the available coverage and evaluation setups do not establish a market leader. BenchLM does not use these results to rank models overall.
GLM-5.2
Z.AI
Artificial Analysis τ²-bench
glm-5-2
GLM-4.7-Flash
Z.AI
Artificial Analysis τ²-bench
glm-4-7-flash
Claude Fable 5
Anthropic
Artificial Analysis τ²-bench
claude-fable-5
Benchmark score table (145 models)
ScoreThe published τ²-bench results snapshot places GLM-5.2 first at 99.1%. The third row is 0.6 points behind. The broader top-10 range is 1.4 points, so many of the published results sit in a relatively narrow band.
145 models have been evaluated on τ²-bench results. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. τ²-bench results is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About τ²-bench results
Year
2025
Tasks
Airline, retail, and telecom customer-service task sets
Format
Published domain success or pass^k results
Difficulty
Dual-control customer-service workflows
τ²-bench extends the original benchmark with a dual-control telecom domain where the agent and simulated user can both act through tools. The maintained framework also includes airline and retail. BenchLM keeps each exact source attached and labels the published setup instead of treating every row as one controlled run.
BenchLM freshness & provenance
Version
τ²-Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
Are all τ²-bench scores directly comparable?
No. Match the domain, task release, agent model, user model, scaffold, prompts, number of trials, and pass^k definition. BenchLM labels each sourced row so a telecom result or third-party implementation is not silently treated as the same setup as an airline, retail, or aggregate result.
Does a high τ²-bench score prove production support reliability?
No. It measures success in simulated customer-service environments under a reported setup. Production identity checks, permission boundaries, changing policies, latency, cost, monitoring, and human escalation still need separate testing.
Compare Top Models on τ²-bench results
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.