Skip to main content

Benchmark profile

Toolathlon

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

Data verified

Benchmark score on Toolathlon — July 23, 2026

BenchLM mirrors the published score view for Toolathlon. Muse Spark 1.1 leads the public snapshot at 75.6% , followed by Claude Opus 4.8 (59.9%) and GPT-5.6 Sol (58%). BenchLM does not use these results to rank models overall.

26 modelsAgenticCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (26 models)

Score
1
Muse Spark 1.1Meta · Closed
75.6%
2
Claude Opus 4.8Anthropic · Closed
59.9%
3
GPT-5.6 SolOpenAI · Closed
58%
4
Gemini 3.5 FlashGoogle · Closed
56.5%
5
GPT-5.5OpenAI · Closed
55.6%
6
GPT-5.4OpenAI · Closed
54.6%
7
GPT-5.6 LunaOpenAI · Closed
53.4%
8
GPT-5.6 TerraOpenAI · Closed
53.1%
9
DeepSeek V4 Pro (Max)DeepSeek · Open weight
51.8%
10
Kimi K2.6Moonshot AI · Open weight
50%
11
Step 3.7 FlashStepFun · Open weight
49.5%
12
DeepSeek V4 Pro (High)DeepSeek · Open weight
49%
13
GLM-5.2Z.AI · Open weight
48.2%
14
DeepSeek V4 Flash (Max)DeepSeek · Open weight
47.8%
15
DeepSeek V4 ProDeepSeek · Open weight
46.3%
16
MiniMax M2.7MiniMax · Open weight
46.3%
17
Claude Opus 4.5Anthropic · Closed
43.5%
18
DeepSeek V4 Flash (High)DeepSeek · Open weight
43.5%
19
GPT-5.4 miniOpenAI · Closed
42.9%
20
DeepSeek V4 FlashDeepSeek · Open weight
40.7%
21
Qwen3.6 PlusAlibaba · Closed
39.8%
22
GLM-5Z.AI · Open weight
38%
23
Qwen3.5 397BAlibaba · Open weight
36.3%
24
GPT-5.4 nanoOpenAI · Closed
35.5%
25
Kimi K2.5Moonshot AI · Open weight
27.8%
26
Qwen3.6-35B-A3BAlibaba · Open weight
26.9%

The published Toolathlon snapshot places Muse Spark 1.1 first at 75.6%. The third row is 17.6 points behind. The broader top-10 range is 25.6 points, so the table still separates the published systems.

26 models have been evaluated on Toolathlon. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Toolathlon is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Toolathlon

Year

2026

Tasks

Multi-tool workflows

Format

Interactive tool-calling evaluation

Difficulty

Advanced tool use

Toolathlon is useful for judging whether a model can do more than answer in chat and instead complete multi-step tool workflows.

BenchLM freshness & provenance

Version

Toolathlon 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Toolathlon measure?

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

Which model scores highest on Toolathlon?

Muse Spark 1.1 by Meta currently leads with a score of 75.6% on Toolathlon.

How many models are evaluated on Toolathlon?

26 AI models have been evaluated on Toolathlon on BenchLM.

Last updated: July 23, 2026 · BenchLM version Toolathlon 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.