Skip to main content
Agentic benchmark report

Best LLMs for AgenticJuly 2026 Leaderboard

As of July 2026, the top agentic model on the BenchLM leaderboard is GPT-5.6 Sol with a BenchAlign agentic score of 75.2.

Data refreshed:

Tool use, browser research, and computer-use workflows

Decision lens: the score determines position; Supported and Estimated labels describe the evidence behind that position without removing sparsely reported models.

Data refreshed
July 23, 2026
Ranked
119 of 290 models
Supported / Estimated
39 / 80
Weighted evidence
3 of 27 benchmarks
27 tracked benchmarks

Terminal-Bench 2.0, BrowseComp, OSWorld-Verified, OSWorld 2.0, CyberGym, Cybench, ExploitGym, JobBench, BrowseComp-VL, OSWorld, AndroidWorld, WebVoyager, MCP Atlas, Toolathlon, Finance Agent v2, GDPval-AA, ZClawBench, Tau2-Telecom, DeepSearchQA, Tau2-Airline, PinchBench, OpenHands Index, SWE-Atlas Refactoring, BFCL v4, MLE-Bench Lite, MM-ClawBench, Gert Labs

Scope: Terminal/tool use, Browser research, Computer use

Best Agentic picks

BenchLM summaries for agentic plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.

How BenchLM scores these

Agentic AI Leaderboard

Primary score: BenchAlign agentic score. Higher values rank first. Use the Show metric control to change the value shown in each row.

Updated Embed leaderboard

Supported positions have diverse direct evidence. Estimated positions remain ranked but carry wider uncertainty.

Filters
Supported positions have diverse direct evidence. Estimated positions remain ranked with wider uncertainty.
Rank / modelWeighted Agentic
1
GPT-5.6 SolOpenAI · ClosedSupported
75.2%
2
GPT-5.6 TerraOpenAI · ClosedSupported
73.1%
3
Muse Spark 1.1Meta · ClosedSupported
68.6%
4
Kimi K3Moonshot AI · ClosedSupported
66.6%
5
Claude Opus 4.7 (Adaptive)Anthropic · ClosedSupported
65.9%
6
GPT-5.5OpenAI · ClosedSupported
65.8%
7
Claude Mythos 5Anthropic · ClosedSupported
65.0%
8
Claude Opus 4.8Anthropic · ClosedSupported
64.0%
9
Claude Fable 5Anthropic · ClosedSupported
63.9%
10
GPT-5.3 CodexOpenAI · ClosedEstimated
61.2%
11
GPT-5.5 ProOpenAI · ClosedEstimated
60.5%
12
Grok 4.5xAI · ClosedSupported
60%
13
Gemini 3 ProGoogle · ClosedEstimated
59.7%
14
Muse SparkMeta · ClosedEstimated
59.4%
15
Claude Sonnet 5Anthropic · ClosedSupported
58.7%
16
GPT-5.6 LunaOpenAI · ClosedSupported
58.5%
17
GPT-5.4 ProOpenAI · ClosedEstimated
58.2%
18
Qwen 3.6 Max (preview)Alibaba · ClosedEstimated
57.2%
19
GPT-5.4OpenAI · ClosedSupported
57%
20
Claude Opus 4.6 (Adaptive)Anthropic · ClosedEstimated
55.4%
21
Claude Opus 4.6Anthropic · ClosedSupported
55.2%
22
Claude Opus 4.7Anthropic · ClosedSupported
54.8%
23
GLM-5Z.AI · Open weightEstimated
54.8%
24
GLM-5.2Z.AI · Open weightSupported
54.6%
25
SWE-1.7Cognition · ClosedEstimated
54.4%

Top AI Models for AgenticJuly 2026

As of July 2026, GPT-5.6 Sol leads the BenchAlign agentic leaderboard with a score of 75.2, followed by GPT-5.6 Terra (73.1) and Muse Spark 1.1 (68.6). BenchLM is currently showing 39 Supported and 80 Estimated models in this category.

What changed

GPT-5.6 Sol ranks #1 at 75.2 with a Supported evidence label.

GPT-5.6 Terra ranks #2 at 73.1 with a Supported evidence label.

Muse Spark 1.1 ranks #3 at 68.6 with a Supported evidence label.

Top models by benchmark

Agentic software engineering and terminal task completion benchmark(38% of category score)

RankModelReported score
2Kimi K3~88.3

Score in Context

What these scores mean

BenchAlign places direct benchmarks and independent external signals on a common calibrated scale. The score is relative to the current evidence universe; it is not a raw percentage from any single test.

Known limitations

Estimated rows have less diverse direct evidence and wider uncertainty. They remain ranked so a newly released model is not treated as weak merely because fewer benchmark publishers have evaluated it.

How we weight

This lens combines category-relevant external evidence with admitted benchmark protocols. Evidence sources are calibrated for difficulty before aggregation, and no generated benchmark row contributes to the score.

Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.

The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.

Scroll horizontally to read the full evidence ledger.

Agentic benchmark weights, ranking status, and descriptions
BenchmarkWeightStatusDescription
Terminal-Bench 2.038%WeightedAgentic software engineering and terminal task completion benchmark
BrowseComp28%WeightedWeb research benchmark for browsing agents
OSWorld-Verified34%WeightedComputer-use benchmark for GUI task completion
OSWorld 2.0Display onlyA long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.
CyberGymDisplay onlyCybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.
CybenchDisplay onlyA cybersecurity benchmark of professional Capture the Flag tasks for measuring autonomous cyber agent capability and risk.
ExploitGymDisplay onlyA controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.
JobBenchDisplay onlyAn occupational agent benchmark for professional workflows that workers say they most want delegated to AI.
BrowseComp-VLDisplay onlyVision-language browsing benchmark for multimodal web research and tool-use tasks.
OSWorldDisplay onlyComputer-use benchmark for GUI task completion across the broader OSWorld task suite.
AndroidWorldDisplay onlyAndroid GUI agent benchmark for task completion across mobile app workflows.
WebVoyagerDisplay onlyBrowser agent benchmark for completing multi-step workflows on live websites.
MCP AtlasDisplay onlyTool-calling benchmark for Model Context Protocol integrations and multi-tool coordination
ToolathlonDisplay onlyGeneral tool-calling benchmark for multi-step API and tool usage
Finance Agent v2Display onlyFinancial analysis and decision-making benchmark for agentic expert tasks.
GDPval-AADisplay onlyReal-world agentic knowledge-work evaluation reported as an Elo score.
ZClawBenchDisplay onlyZ.AI's OpenClaw workflow benchmark for broad agent tasks across research, office work, data analysis, devops, automation, and security.
Tau2-TelecomDisplay onlyTelecom-focused tool-use benchmark for structured API workflows
DeepSearchQADisplay onlyAgentic browsing benchmark for list-style question answering with browser tools.
Tau2-AirlineDisplay onlyAirline-domain tool-use benchmark for structured workflow execution and API correctness.
PinchBenchDisplay onlyAn OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.
OpenHands IndexDisplay onlyA holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.
SWE-Atlas RefactoringDisplay onlyA Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
BFCL v4Display onlyFunction-calling benchmark for tool selection, schema adherence, and argument correctness.
MLE-Bench LiteDisplay onlyA lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings.
MM-ClawBenchDisplay onlyAn OpenClaw-derived agent benchmark covering practical work and life tasks such as office document delivery, research, planning, and code maintenance.
Gert LabsDisplay onlyComposite game-environment leaderboard score across Gert Labs agentic coding, one-shot coding, and social decision-making modes.

About Agentic Benchmarks

Agentic software engineering and terminal task completion benchmark

Common questions

What is an agentic LLM benchmark?

Agentic benchmarks evaluate whether AI models can complete multi-step workflows using tools, browsers, terminals, or software interfaces instead of only answering in chat.

Which benchmarks matter for AI agents?

Key agentic benchmarks include Terminal-Bench 2.0 for terminal tasks, BrowseComp for web research, and OSWorld-Verified for computer-use workflows.

Why do agentic benchmarks matter in 2026?

Agentic benchmarks matter because many modern products rely on models that can browse, plan, use tools, and complete end-to-end tasks rather than only generate text.

Agentic benchmark updates

Agentic is the fastest-moving category. Don't fall behind.

One email each week. Unsubscribe anytime.

Related