Benchmark profile
τ³-Bench Tool-Agent-User Evaluation (τ³-bench results)
τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.
Data verifiedHow to read this leaderboard
Editorial review by Glevd · 2026-07-15
Use each result with its attached source and setup label. Match domain or published average, task release, modality, agent and user models, scaffold, prompts, trial count, and pass^k policy before comparison. The sorted table does not imply one controlled benchmark run.
Operator receipt: 10 sourced rows are currently displayable on this page; the highest published score among these 10 models is Mistral Medium 3.5 128B at 91.4%, which does not establish a market leader.
Honest limit: Ten sourced rows come from several provider reports, and their labels include telecom and broader published averages. BenchLM did not rerun them. The framework's active task fixes and expanding modalities also mean an older result may not describe the current release.
Benchmark score on τ³-bench sourced results — July 23, 2026
BenchLM mirrors the published score view for τ³-bench sourced results. The public snapshot contains 10 models. Mistral Medium 3.5 128B has the highest published score at 91.4%, but the available coverage and evaluation setups do not establish a market leader. BenchLM does not use these results to rank models overall.
Mistral Medium 3.5 128B
Mistral
τ³-bench Telecom
mistral-medium-3-5-128b
MiMo-V2.5-Pro
Xiaomi
τ³-bench published setup
mimo-v2-5-pro
Nemotron 3 Ultra
NVIDIA
τ³-bench published average
nemotron-3-ultra-500b
Benchmark score table (10 models)
ScoreThe published τ³-bench results snapshot places Mistral Medium 3.5 128B first at 91.4%. The third row is 20.5 points behind. The broader top-10 range is 25.8 points, so the table still separates the published systems.
10 models have been evaluated on τ³-bench results. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. τ³-bench results is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About τ³-bench results
Year
2026
Tasks
Corrected customer-service tasks plus knowledge and voice evaluation modes
Format
Published domain or average success results
Difficulty
Long-horizon, multimodal, and knowledge-aware tool use
The maintained repository now identifies the framework as τ³-bench and documents text, voice, telecom, airline, retail, and knowledge-aware banking evaluation. BenchLM keeps provider-published τ³ rows separate from original TAU-bench and τ²-bench rows.
BenchLM freshness & provenance
Version
τ³-bench results 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
Is τ³-bench the same as original TAU-bench?
No. It is the maintained successor framework with corrected tasks and newer domains and modalities. BenchLM keeps original TAU-bench, τ²-bench, τ² Airline, and τ³-bench on separate routes.
Can every τ³-bench result be ranked together?
No. Match the domain or average, release, modality, agent and user models, scaffold, prompts, trials, and pass^k policy. Provider tables without those same controls are useful source receipts, not one apples-to-apples leaderboard.
Compare Top Models on τ³-bench results
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.