Benchmark profile
τ²-Bench Airline Domain (τ²-bench Airline)
τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.
Data verifiedHow to read this leaderboard
Editorial review by Glevd · 2026-07-15
Treat the current display as one source receipt, not a market ranking. Compare another airline result only when the task release, agent and user models, scaffold, prompts, trial count, and avg@1 or pass^k policy match.
Operator receipt: 1 sourced row is currently displayable on this page; the only currently displayable sourced row is ZAYA1-74B-Preview at 56.1%.
Honest limit: Only one of six stored airline rows currently has an exact source attachment. The other five remain withheld. One sourced row cannot establish a leader, and a simulated airline workflow does not cover production authentication, permissions, live inventory, policy drift, or human escalation.
Benchmark score on τ²-bench Airline — July 23, 2026
BenchLM mirrors the published score view for τ²-bench Airline. ZAYA1-74B-Preview is the only currently displayable sourced row, at 56.1%. BenchLM does not use these results to rank models overall.
Benchmark score table (1 model)
ScoreAbout τ²-bench Airline
Year
2025
Tasks
Airline customer-service tasks
Format
Domain success under a published trial policy
Difficulty
Policy-constrained airline support workflows
This lane owns results explicitly labeled for the airline domain. It stays separate from telecom, retail, aggregates, the archived original TAU-bench release, and newer τ³-bench runs.
BenchLM freshness & provenance
Version
τ²-bench Airline 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does τ²-bench Airline measure?
τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.
Which model scores highest on τ²-bench Airline?
ZAYA1-74B-Preview by Zyphra currently leads with a score of 56.1% on τ²-bench Airline.
How many models are evaluated on τ²-bench Airline?
1 AI models have been evaluated on τ²-bench Airline on BenchLM.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.