Skip to main content

Benchmark profile

OSWorld-Verified

OSWorld-Verified is the July 2025 repaired release of OSWorld's real-computer evaluation. It measures whether a model-agent system can finish desktop and web tasks from configured starting states, with success checked by execution-based evaluators.

Data verified

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Read this table as a ledger of sourced model-agent results. Match the task count, environment revision, model and agent variant, general model versus specialized model versus agentic framework, screenshot or accessibility-tree observation, action interface, prompt and scaffold, action budget, and attempt policy before comparing scores. The benchmark-hosted table uses unified settings; provider-reported rows are not automatically part of that controlled comparison.

Operator receipt: 27 sourced rows are currently displayable on this page; the leading published row is Claude Mythos 5 at 85%.

Honest limit: A success rate reflects the complete model-agent system, not the base model alone. Fixed applications and configured starting states do not reproduce every production login, permission boundary, network failure, application update, latency, cost, or safety control. Eight Google Drive tasks may be excluded under the official policy, and rows without displayable exact-source records remain withheld. OSWorld 2.0 results are not interchangeable with OSWorld-Verified results.

Top models on OSWorld-Verified — July 23, 2026

As of July 23, 2026, Claude Mythos 5 leads the OSWorld-Verified leaderboard with 85% , followed by Claude Fable 5 (85%) and Claude Opus 4.8 (83.4%).

27 modelsAgentic34% of category scoreCurrentUpdated July 23, 2026

Leaderboard (27 models)

Score
1
Claude Mythos 5Anthropic · Closed
85%
2
Claude Fable 5Anthropic · Closed
85%
3
Claude Opus 4.8Anthropic · Closed
83.4%
4
Gemini 3.6 FlashGoogle · Closed
83%
5
Holo3-35B-A3BH Company · Open weight
82.6%
6
Claude Sonnet 5Anthropic · Closed
81.2%
7
Muse Spark 1.1Meta · Closed
80.8%
8
Holo3-122B-A10BH Company · Closed
78.8%
9
GPT-5.5OpenAI · Closed
78.7%
10
Gemini 3.5 FlashGoogle · Closed
78.4%
11
Claude Opus 4.7 (Adaptive)Anthropic · Closed
78%
12
GPT-5.4OpenAI · Closed
75%
13
Gemini 3.5 Flash-LiteGoogle · Closed
74%
14
Qwen3.7 PlusAlibaba · Closed
73.3%
15
Kimi K2.6Moonshot AI · Open weight
73.1%
16
Claude Opus 4.6Anthropic · Closed
72.7%
17
Claude Sonnet 4.6Anthropic · Closed
72.1%
18
GPT-5.4 miniOpenAI · Closed
72.1%
19
MiniMax M3MiniMax · Open weight
70.1%
20
Claude Opus 4.5Anthropic · Closed
66.3%
21
GPT-5.3 CodexOpenAI · Closed
64.7%
22
Claude Sonnet 4.5Anthropic · Closed
61.4%
23
Qwen3.5-122B-A10BAlibaba · Open weight
58%
24
Qwen3.5-27BAlibaba · Open weight
56.2%
25
Qwen3.5-35B-A3BAlibaba · Open weight
54.5%
26
GPT-5.2OpenAI · Closed
47.3%
27
GPT-5.4 nanoOpenAI · Closed
39%

According to BenchLM.ai, Claude Mythos 5 leads the OSWorld-Verified benchmark with a score of 85%, followed by Claude Fable 5 (85%) and Claude Opus 4.8 (83.4%). The top models are clustered within 1.6 points, suggesting this benchmark is nearing saturation for frontier models.

27 models have been evaluated on OSWorld-Verified. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Within that category, OSWorld-Verified contributes 34% of the category score, so strong performance here directly affects a model's overall ranking.

About OSWorld-Verified

Year

2025

Tasks

369 real-world computer tasks (361 when eight Google Drive tasks are excluded)

Format

Execution-based interactive task success

Difficulty

Multi-step desktop and cross-application workflows

The release covers 369 real-world tasks across desktop and web applications. The maintainers allow eight Google Drive tasks to be manually configured or excluded, making a 361-task run officially acceptable. Public evaluation requires the maintainers to run the agent or review monitoring data and trajectories. OSWorld 2.0 is a newer, separate protocol.

BenchLM freshness & provenance

Version

OSWorld Verified

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does OSWorld-Verified measure?

OSWorld-Verified measures whether a computer-use agent can complete 369 tasks across real desktop and web applications. Each task starts from a configured state and is checked with an execution-based evaluator. Eight Google Drive tasks may be manually configured or excluded, producing an officially accepted 361-task run.

Are OSWorld-Verified scores directly comparable?

Only when the setup matches. Compare the same task set, environment revision, model-agent variant, observation and action interface, prompt or scaffold, action budget, and attempt policy. The official leaderboard separates general models, specialized models, and agentic frameworks; provider-published scores may use different conditions.

Can OSWorld-Verified choose the best computer-use model?

Not alone. It is strong evidence for desktop task completion, but fixed applications cannot reproduce every production login, permission, network, app-version, latency, cost, or safety condition. Use matched OSWorld-Verified results with workflow trials, and treat OSWorld 2.0 as a separate newer protocol rather than interchangeable evidence.

Last updated: July 23, 2026 · BenchLM version OSWorld Verified

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.