Skip to main content

Benchmark profile

Claw-Eval

A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.

Data verified

How BenchLM shows Claw-Eval

BenchLM mirrors the official Claw-Eval 2026-05-09 leaderboard snapshot. The source benchmark contains 300 human-verified tasks, 2,159 rubric items, and uses Pass^3 as the primary metric across 3 independent trials.

The public Claw-Eval site separates the 199-task general plus multi-turn agent table from the 101-task native multimodal table. BenchLM sorts this page by the primary general plus multi-turn Pass^3 table and preserves native multimodal split scores in the mirrored snapshot metadata.

Claw-Eval is display only on BenchLM. It is strong evidence about agent reliability, but the public rows are benchmark-harness results rather than normalized model-only rankings, so they are excluded from BenchLM overall and category scores.

26 primary agent rows11 native multimodal rows300 tasks2,159 rubricsDisplay only

Pass^3 on Claw-Eval — 2026-05-09 snapshot

BenchLM mirrors the published pass^3 view for Claw-Eval. Claude Opus 4.6 leads the public snapshot at 70.4% , followed by Step 3.7 Flash (68.3%) and Claude Sonnet 4.6 (67.8%). BenchLM does not use these results to rank models overall.

26 modelsAgenticCurrentDisplay onlyUpdated 2026-05-09 snapshot

Pass^3 table (26 models)

Score
1
Claude Opus 4.6Anthropic · Closed
70.4%
2
Step 3.7 FlashStepFun · Open weight
68.3%
3
Claude Sonnet 4.6Anthropic · Closed
67.8%
4
MiMo-V2.5-ProXiaomi · Closed
63.8%
5
Muse SparkMeta · Closed
63.8%
6
Kimi K2.6Moonshot AI · Open weight
62.3%
7
MiMo-V2.5Xiaomi · Closed
62.3%
8
GLM-5.1Z.AI · Open weight
62.3%
9
GPT-5.4OpenAI · Closed
60.3%
10
DeepSeek V4 ProDeepSeek · Open weight
59.8%
11
58.8%
12
Qwen3.6 PlusAlibaba · Closed
58.8%
13
Gemini 3.1 ProGoogle · Closed
57.8%
14
DeepSeek V4 FlashDeepSeek · Open weight
57.8%
15
MiMo-V2-ProXiaomi · Closed
57.8%
16
Qwen3.5 397BAlibaba · Open weight
56.8%
17
GLM-5-TurboZ.AI · Closed
55.8%
18
GLM-5V-TurboZ.AI · Closed
53.8%
19
Kimi K2.5Moonshot AI · Open weight
52.3%
20
51.8%
21
Gemini 3 FlashGoogle · Closed
49.2%
23
MiniMax M2.7MiniMax · Open weight
48.7%
24
MiMo-V2-OmniXiaomi · Closed
45.2%
25
DeepSeek V3.2DeepSeek · Open weight
40.2%
26
Nemotron 3 Super 100BNVIDIA · Open weight
5.5%

The published Claw-Eval snapshot places Claude Opus 4.6 first at 70.4%. The third row is 2.6 points behind. The broader top-10 range is 10.6 points, so the table still separates the published systems.

26 models have been evaluated on Claw-Eval. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Claw-Eval is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Claw-Eval

Year

2026

Tasks

300 tasks, 2,159 rubrics

Format

End-to-end autonomous-agent evaluation with Pass^3 scoring

Difficulty

Real-world general, multi-turn, and native multimodal agent execution

Claw-Eval v1.1.0 evaluates autonomous agents on full-trajectory tasks audited for completion, safety, and robustness. Its primary Pass^3 metric requires a task to pass in all three independent trials, reducing lucky-run effects. BenchLM mirrors the official leaderboard as display-only because rows reflect benchmark harness execution as well as model capability.

BenchLM freshness & provenance

Version

Claw-Eval 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Claw-Eval measure?

A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.

Which model leads the published Claw-Eval snapshot?

Claude Opus 4.6 currently leads the published Claw-Eval snapshot with 70.4% pass^3. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Claw-Eval?

26 AI models are included in BenchLM's mirrored Claw-Eval snapshot, based on the public leaderboard captured on 2026-05-09 snapshot.

Last updated: 2026-05-09 snapshot · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.