Benchmark profile
OpenHands Index
A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.
How BenchLM shows OpenHands Index
BenchLM mirrors the official OpenHands Index REST API snapshot from July 21, 2026 snapshot. The source evaluates 33 coding-agent model variants across 5 categories: Issue Resolution, Frontend, Greenfield, Testing, Information Gathering.
OpenHands Index is display only on BenchLM. It is a valuable agentic software-engineering reference, but its rows combine model, SDK version, agent harness, cost, runtime, and per-benchmark result links, so BenchLM keeps it separate from weighted model-only rankings.
Average agent score on OpenHands Index — July 21, 2026 snapshot
BenchLM mirrors the published average agent score view for OpenHands Index. Claude Fable 5 leads the public snapshot at 81.0% , followed by Claude Opus 4.8 (71.9%) and Claude Opus 4.7 (Adaptive) (69.7%). BenchLM does not use these results to rank models overall.
Claude Fable 5
Anthropic
claude-fable-5
Claude Opus 4.8
Anthropic
claude-opus-4-8
Claude Opus 4.7 (Adaptive)
Anthropic
claude-opus-4-7
Average agent score table (33 models)
ScoreThe published OpenHands Index snapshot places Claude Fable 5 first at 81.0%. The third row is 11.3 points behind. The broader top-10 range is 22.2 points, so the table still separates the published systems.
33 models have been evaluated on OpenHands Index. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. OpenHands Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About OpenHands Index
Year
2025
Tasks
SWE-bench Verified, SWE-bench Multimodal, Commit0, SWT-bench Verified, and GAIA
Format
Macro-average across five coding-agent categories
Difficulty
Real-world software engineering agent tasks
BenchLM mirrors the official OpenHands Index REST API as a display-only agentic software-engineering benchmark. The source reports average agent score, cost, runtime, per-category scores, logs, and visualizations for each model and SDK version.
BenchLM freshness & provenance
Version
OpenHands Index 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does OpenHands Index measure?
A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.
Which model leads the published OpenHands Index snapshot?
Claude Fable 5 currently leads the published OpenHands Index snapshot with 81.0% average agent score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on OpenHands Index?
33 AI models are included in BenchLM's mirrored OpenHands Index snapshot, based on the public leaderboard captured on July 21, 2026 snapshot.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.