Skip to main content

Agent & Tool-Use Benchmarks

Which AI models handle function calling, MCP tool use, browsing, and multi-step agent workflows best? Verified-ranked results across 24 agentic benchmarks.

Agentic carries 22% weightin BenchLM.ai's overall score — the single biggest category.

This page now shows only core agentic benchmark rows with an attached exact source record. Source-unverified manual rows are excluded from the displayed agentic score and table cells.

Best Agentic Model

GPT-5.6 Sol

Verified score: 92 · OpenAI

Best Open-Weight Agent

Holo3-35B-A3B

Verified score: 82.6 · H Company

Benchmarks Tracked

26 benchmarks

Terminal, browsing, tool-use, and computer-use

Benchmark Categories

Core Weighted (3)

These 3 benchmarks determine agentic rankings

Tool Calling & MCP (6)

Function calling, MCP tool use, and structured workflows

Computer & Browser Use (5)

Desktop GUI, mobile, and browser navigation tasks

Specialized (4)

Domain-specific agentic tasks across ML, research, and airline

Top 15 Models by Weighted Agentic Score

OpenAIAnthropicGoogleMetaDeepSeekMistralxAIAlibaba

74 models match these filters.

CSVJSON
RankModelVerified score
1
GPT-5.6 SolOpenAI · proprietaryTB 91.9BC 92.2OS
92agentic / 100
2
GPT-5.5 ProOpenAI · proprietaryTB BC 90.1OS
90.1agentic / 100
3
Kimi K3Moonshot AI · proprietaryTB 88.3BC 91.2OS
89.5agentic / 100
4
GPT-5.4 ProOpenAI · proprietaryTB BC 89.3OS
89.3agentic / 100
5
GPT-5.6 TerraOpenAI · proprietaryTB 87.4BC 87.5OS
87.4agentic / 100
7
Claude Fable 5Anthropic · proprietaryTB 84.3BC OS 85
84.6agentic / 100
8
GPT-5.6 LunaOpenAI · proprietaryTB 84.7BC 83.3OS
84.1agentic / 100
9
Grok 4.5xAI · proprietaryTB 83.3BC OS
83.3agentic / 100
11
Holo3-35B-A3BH Company · open weightTB BC OS 82.56
82.6agentic / 100
13
Claude Sonnet 5Anthropic · proprietaryTB 80.4BC 84.7OS 81.2
81.9agentic / 100
14
GPT-5.5OpenAI · proprietaryTB 82BC 84.4OS 78.7
81.6agentic / 100
15
SWE-1.7Cognition · proprietaryTB 81.5BC OS
81.5agentic / 100
16
GLM-5.2Z.AI · open weightTB 81BC OS
81agentic / 100
17
Muse Spark 1.1Meta · proprietaryTB 80BC OS 80.8
80.4agentic / 100
18
Claude Opus 4.8Anthropic · proprietaryTB 74.6BC 84.3OS 83.4
80.3agentic / 100
19
Sakana FuguSakana AI · proprietaryTB 80.2BC OS
80.2agentic / 100
20
Holo3-122B-A10BH Company · proprietaryTB BC OS 78.85
78.9agentic / 100
21
Ornith-1.0-397BDeepReinforce AI · open weightTB 77.5BC OS
77.5agentic / 100
23
GPT-5.4OpenAI · proprietaryTB 75.1BC 82.7OS 75
77.2agentic / 100
24
Agents-A1InternScience · open weightTB BC 75.51OS
75.5agentic / 100
27
Kimi K2.6Moonshot AI · open weightTB 66.7BC 83.2OS 73.1
73.5agentic / 100
28
Claude Opus 4.6Anthropic · proprietaryTB 65.4BC 83.7OS 72.7
73agentic / 100
29
MiniMax M3MiniMax · open weightTB 66BC 83.52OS 70.06
72.3agentic / 100
30
Qwen3.7 PlusAlibaba · proprietaryTB 70.3BC OS 73.3
71.7agentic / 100
31
GPT-5.3 CodexOpenAI · proprietaryTB 77.3BC OS 64.7
71.4agentic / 100
33
Laguna S 2.1Poolside · open weightTB 70.2BC OS
70.2agentic / 100
34
Qwen3.7 MaxAlibaba · proprietaryTB 69.7BC OS
69.7agentic / 100
35
InklingThinking Machines Lab · open weightTB 63.8BC 77.1OS
69.4agentic / 100
36
Composer 2.5Cursor · proprietaryTB 69.3BC OS
69.3agentic / 100
37
MiMo-V2.5-ProXiaomi · proprietaryTB 68.4BC OS
68.4agentic / 100
38
Step 3.7 FlashStepFun · open weightTB 59.5BC 75.82OS
66.4agentic / 100
39
MiMo-V2.5Xiaomi · proprietaryTB 65.8BC OS
65.8agentic / 100
40
GPT-5.4 miniOpenAI · proprietaryTB 60BC OS 72.1
65.7agentic / 100
41
GLM-5.1Z.AI · open weightTB 63.5BC 68OS
65.4agentic / 100
44
Ornith-1.0-35BDeepReinforce AI · open weightTB 64.2BC OS
64.2agentic / 100
47
Claude Opus 4.5Anthropic · proprietaryTB 59.3BC OS 66.3
62.6agentic / 100
48
Composer 2Cursor · proprietaryTB 61.7BC OS
61.7agentic / 100
49
Qwen3.6 PlusAlibaba · proprietaryTB 61.6BC OS
61.6agentic / 100
50
Qwen3.6-27BAlibaba · open weightTB 59.3BC OS
59.3agentic / 100
51
DeepSeek V4 ProDeepSeek · open weightTB 59.1BC OS
59.1agentic / 100
52
Muse SparkMeta · proprietaryTB 59BC OS
59agentic / 100
53
MiniMax M2.7MiniMax · open weightTB 57BC OS
57agentic / 100
54
Qwen3.5 397BAlibaba · open weightTB 52.5BC 62OS
56.5agentic / 100
56
GLM-5Z.AI · open weightTB 56.2BC OS
56.2agentic / 100
57
GPT-5.2OpenAI · proprietaryTB BC 65.8OS 47.3
55.7agentic / 100
60
Kimi K2.5Moonshot AI · open weightTB 50.8BC 60.6OS
55agentic / 100
62
Hy3 PreviewTencent · open weightTB 54.4BC OS
54.4agentic / 100
63
Qwen3.5-27BAlibaba · open weightTB 41.6BC 61OS 56.2
52agentic / 100
64
Qwen3.6-35B-A3BAlibaba · open weightTB 51.5BC OS
51.5agentic / 100
68
Grok 4.20xAI · proprietaryTB 47.1BC OS
47.1agentic / 100
69
MAI-Thinking-1Microsoft · proprietaryTB 46BC OS
46agentic / 100
70
Laguna M.1Poolside · proprietaryTB 45.8BC OS
45.8agentic / 100
71
GLM-4.7Z.AI · open weightTB 41BC 52OS
45.7agentic / 100
72
Ornith-1.0-9BDeepReinforce AI · open weightTB 43.1BC OS
43.1agentic / 100
73
GPT-5.4 nanoOpenAI · proprietaryTB 46.3BC OS 39
42.9agentic / 100
74
Laguna XS.2Poolside · open weightTB 35.7BC OS
35.7agentic / 100

Agentic score = weighted average of Terminal-Bench 2.0 (40%), OSWorld-Verified (35%), and BrowseComp (25%), normalized by available weights. This page intentionally stays on BenchLM's verified ranking lane and only includes exact-source rows. Display-only benchmarks (MCP Atlas, Toolathlon, etc.) are tracked but do not affect rankings.

Frequently Asked Questions

What are LLM agent benchmarks?

Agent benchmarks test whether AI models can go beyond answering questions and actually complete multi-step tasks: browsing the web, writing and running code in a terminal, calling external APIs via function calling, and operating desktop or mobile interfaces. They measure real-world usefulness for autonomous workflows.

What is function calling and why does it matter?

Function calling (or tool use) lets an LLM invoke external tools, APIs, or databases as part of its response. This is critical for building AI agents that can search the web, query databases, send emails, or control other software. Benchmarks like BFCL v4 and Toolathlon specifically measure how reliably models select the right function and pass correct arguments.

What is MCP (Model Context Protocol)?

MCP is an open standard for connecting LLMs to external tools and data sources. MCP Atlas and MCP-Tasks benchmark how well models work with MCP-backed integrations. Strong MCP performance means a model integrates well into tool-rich agent architectures.

Why does agentic carry the most weight in BenchLM scores?

Agentic carries 22% of BenchLM's overall score because the ability to use tools, browse, and complete multi-step tasks is the strongest differentiator between models in production use. A model that scores well on knowledge but cannot reliably call functions or navigate software has limited real-world utility for agent workflows.

Which models are best for building AI agents?

Currently, GPT-5.6 Sol by OpenAI leads BenchLM's verified agentic rankings with a score of 92. The best open-weight agent model is Holo3-35B-A3B (82.6). Check the leaderboard above for the full verified ranking.

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.