Skip to main content

Benchmark profile

OpenHands Index

A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.

How BenchLM shows OpenHands Index

BenchLM mirrors the official OpenHands Index REST API snapshot from July 21, 2026 snapshot. The source evaluates 33 coding-agent model variants across 5 categories: Issue Resolution, Frontend, Greenfield, Testing, Information Gathering.

OpenHands Index is display only on BenchLM. It is a valuable agentic software-engineering reference, but its rows combine model, SDK version, agent harness, cost, runtime, and per-benchmark result links, so BenchLM keeps it separate from weighted model-only rankings.

33 model variants17 open models16 closed models5 benchmarksDisplay only

Average agent score on OpenHands Index — July 21, 2026 snapshot

BenchLM mirrors the published average agent score view for OpenHands Index. Claude Fable 5 leads the public snapshot at 81.0% , followed by Claude Opus 4.8 (71.9%) and Claude Opus 4.7 (Adaptive) (69.7%). BenchLM does not use these results to rank models overall.

33 modelsAgenticCurrentDisplay onlyUpdated July 21, 2026 snapshot

Average agent score table (33 models)

Score
1
Claude Fable 5Anthropic · Closed
81.0%
2
Claude Opus 4.8Anthropic · Closed
71.9%
3
Claude Opus 4.7 (Adaptive)Anthropic · Closed
69.7%
4
Claude Opus 4.6Anthropic · Closed
66.7%
5
GPT-5.5OpenAI · Closed
65.9%
6
GPT-5.4OpenAI · Closed
64.3%
7
Gemini 3.5 FlashGoogle · Closed
62.6%
8
Claude Opus 4.5Anthropic · Closed
60.6%
9
Gemini 3.1 ProGoogle · Closed
60.6%
10
GPT-5.2OpenAI · Closed
58.8%
11
GPT-5.2-CodexOpenAI · Closed
58.3%
12
GLM-5.1Z.AI · Open weight
58.2%
13
MiniMax M3MiniMax · Open weight
57.2%
14
Kimi K2.6Moonshot AI · Open weight
57.1%
15
Claude Sonnet 4.5Anthropic · Closed
53.0%
16
Qwen3.6 PlusAlibaba · Closed
52.9%
17
GLM-5Z.AI · Open weight
49.4%
18
Kimi K2.5Moonshot AI · Open weight
49.2%
19
Gemini 3 ProGoogle · Closed
49.0%
20
Gemini 3 FlashGoogle · Closed
49.0%
21
DeepSeek V3.2 (Thinking)DeepSeek · Open weight
45.7%
22
MiniMax M2.5MiniMax · Closed
45.2%
23
Claude Sonnet 4.6Anthropic · Closed
44.5%
24
43.8%
25
MiniMax M2.7MiniMax · Open weight
43.4%
26
GLM-4.7Z.AI · Open weight
42.3%
27
41.2%
28
Kimi K2.5 (Reasoning)Moonshot AI · Closed
41.0%
29
DeepSeek V4 ProDeepSeek · Open weight
40.7%
30
Qwen3.5 FlashAlibaba · Closed
38.1%
31
Nemotron 3 Super 120B A12BNVIDIA · Open weight
36.2%
32
Trinity-Large-ThinkingArcee AI · Open weight
32.1%
33
30.9%

The published OpenHands Index snapshot places Claude Fable 5 first at 81.0%. The third row is 11.3 points behind. The broader top-10 range is 22.2 points, so the table still separates the published systems.

33 models have been evaluated on OpenHands Index. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. OpenHands Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About OpenHands Index

Year

2025

Tasks

SWE-bench Verified, SWE-bench Multimodal, Commit0, SWT-bench Verified, and GAIA

Format

Macro-average across five coding-agent categories

Difficulty

Real-world software engineering agent tasks

BenchLM mirrors the official OpenHands Index REST API as a display-only agentic software-engineering benchmark. The source reports average agent score, cost, runtime, per-category scores, logs, and visualizations for each model and SDK version.

BenchLM freshness & provenance

Version

OpenHands Index 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does OpenHands Index measure?

A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.

Which model leads the published OpenHands Index snapshot?

Claude Fable 5 currently leads the published OpenHands Index snapshot with 81.0% average agent score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on OpenHands Index?

33 AI models are included in BenchLM's mirrored OpenHands Index snapshot, based on the public leaderboard captured on July 21, 2026 snapshot.

Last updated: July 21, 2026 snapshot · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.