Skip to main content

Benchmark profile

τ³-Bench Tool-Agent-User Evaluation (τ³-bench results)

τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.

Data verified

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use each result with its attached source and setup label. Match domain or published average, task release, modality, agent and user models, scaffold, prompts, trial count, and pass^k policy before comparison. The sorted table does not imply one controlled benchmark run.

Operator receipt: 10 sourced rows are currently displayable on this page; the highest published score among these 10 models is Mistral Medium 3.5 128B at 91.4%, which does not establish a market leader.

Honest limit: Ten sourced rows come from several provider reports, and their labels include telecom and broader published averages. BenchLM did not rerun them. The framework's active task fixes and expanding modalities also mean an older result may not describe the current release.

Benchmark score on τ³-bench sourced results — July 23, 2026

BenchLM mirrors the published score view for τ³-bench sourced results. The public snapshot contains 10 models. Mistral Medium 3.5 128B has the highest published score at 91.4%, but the available coverage and evaluation setups do not establish a market leader. BenchLM does not use these results to rank models overall.

10 modelsAgenticCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (10 models)

Score
1
Mistral Medium 3.5 128BMistral · Open weightτ³-bench Telecom
91.4%
2
MiMo-V2.5-ProXiaomi · Closedτ³-bench published setup
72.9%
3
Nemotron 3 UltraNVIDIA · Open weightτ³-bench published average
70.9%
4
Qwen3.6 PlusAlibaba · Closedτ³-bench published setup
70.7%
5
GLM-5.1Z.AI · Open weightτ³-bench published setup
70.6%
6
Claude Opus 4.5Anthropic · Closedτ³-bench published setup
70.2%
7
Qwen3.5 397BAlibaba · Open weightτ³-bench published setup
68.4%
8
Qwen3.6-35B-A3BAlibaba · Open weightτ³-bench published setup
67.2%
9
Kimi K2.5Moonshot AI · Open weightτ³-bench published setup
65.7%
10
GLM-5Z.AI · Open weightτ³-bench published setup
65.6%

The published τ³-bench results snapshot places Mistral Medium 3.5 128B first at 91.4%. The third row is 20.5 points behind. The broader top-10 range is 25.8 points, so the table still separates the published systems.

10 models have been evaluated on τ³-bench results. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. τ³-bench results is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About τ³-bench results

Year

2026

Tasks

Corrected customer-service tasks plus knowledge and voice evaluation modes

Format

Published domain or average success results

Difficulty

Long-horizon, multimodal, and knowledge-aware tool use

The maintained repository now identifies the framework as τ³-bench and documents text, voice, telecom, airline, retail, and knowledge-aware banking evaluation. BenchLM keeps provider-published τ³ rows separate from original TAU-bench and τ²-bench rows.

BenchLM freshness & provenance

Version

τ³-bench results 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

Is τ³-bench the same as original TAU-bench?

No. It is the maintained successor framework with corrected tasks and newer domains and modalities. BenchLM keeps original TAU-bench, τ²-bench, τ² Airline, and τ³-bench on separate routes.

Can every τ³-bench result be ranked together?

No. Match the domain or average, release, modality, agent and user models, scaffold, prompts, trials, and pass^k policy. Provider tables without those same controls are useful source receipts, not one apples-to-apples leaderboard.

Last updated: July 23, 2026 · BenchLM version τ³-bench results 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.