Skip to main content

Benchmark profile

Toolathlon-Verified

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Data verified

Benchmark score on Toolathlon-Verified — July 23, 2026

BenchLM mirrors the published score view for Toolathlon-Verified. Kimi K3 leads the public snapshot at 73.2% , followed by Laguna S 2.1 (49.7%). BenchLM does not use these results to rank models overall.

2 modelsAgenticCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (2 models)

Score
1
Kimi K3Moonshot AI · Closed
73.2%
2
Laguna S 2.1Poolside · Open weight
49.7%

About Toolathlon-Verified

Year

2026

Tasks

Verified multi-tool workflows

Format

Interactive tool-use score

Difficulty

Advanced tool use

BenchLM keeps Toolathlon-Verified separate from the broader Toolathlon key so provider launch values do not collapse distinct benchmark variants.

BenchLM freshness & provenance

Version

Toolathlon-Verified 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Toolathlon-Verified measure?

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Which model scores highest on Toolathlon-Verified?

Kimi K3 by Moonshot AI currently leads with a score of 73.2% on Toolathlon-Verified.

How many models are evaluated on Toolathlon-Verified?

2 AI models have been evaluated on Toolathlon-Verified on BenchLM.

Compare Top Models on Toolathlon-Verified

Last updated: July 23, 2026 · BenchLM version Toolathlon-Verified 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.