Skip to main content

Benchmark profile

τ²-Bench Airline Domain (τ²-bench Airline)

τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.

Data verified

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Treat the current display as one source receipt, not a market ranking. Compare another airline result only when the task release, agent and user models, scaffold, prompts, trial count, and avg@1 or pass^k policy match.

Operator receipt: 1 sourced row is currently displayable on this page; the only currently displayable sourced row is ZAYA1-74B-Preview at 56.1%.

Honest limit: Only one of six stored airline rows currently has an exact source attachment. The other five remain withheld. One sourced row cannot establish a leader, and a simulated airline workflow does not cover production authentication, permissions, live inventory, policy drift, or human escalation.

Benchmark score on τ²-bench Airline — July 23, 2026

BenchLM mirrors the published score view for τ²-bench Airline. ZAYA1-74B-Preview is the only currently displayable sourced row, at 56.1%. BenchLM does not use these results to rank models overall.

1 modelAgenticCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (1 model)

Score
1
ZAYA1-74B-PreviewZyphra · Open weightτ²-bench Airline · avg@1
56.1%

About τ²-bench Airline

Year

2025

Tasks

Airline customer-service tasks

Format

Domain success under a published trial policy

Difficulty

Policy-constrained airline support workflows

This lane owns results explicitly labeled for the airline domain. It stays separate from telecom, retail, aggregates, the archived original TAU-bench release, and newer τ³-bench runs.

BenchLM freshness & provenance

Version

τ²-bench Airline 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does τ²-bench Airline measure?

τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.

Which model scores highest on τ²-bench Airline?

ZAYA1-74B-Preview by Zyphra currently leads with a score of 56.1% on τ²-bench Airline.

How many models are evaluated on τ²-bench Airline?

1 AI models have been evaluated on τ²-bench Airline on BenchLM.

Last updated: July 23, 2026 · BenchLM version τ²-bench Airline 2025

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.