Benchmark profile
OSWorld 2.0
A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.
Data verifiedBenchmark score on OSWorld 2.0 — July 23, 2026
BenchLM mirrors the published score view for OSWorld 2.0. GPT-5.6 Sol leads the public snapshot at 62.6% , followed by GPT-5.6 Terra (50.2%) and GPT-5.6 Luna (45.6%). BenchLM does not use these results to rank models overall.
GPT-5.6 Sol
OpenAI
gpt-5-6-sol
GPT-5.6 Terra
OpenAI
gpt-5-6-terra
GPT-5.6 Luna
OpenAI
gpt-5-6-luna
Benchmark score table (12 models)
ScoreThe published OSWorld 2.0 snapshot places GPT-5.6 Sol first at 62.6%. The third row is 17.0 points behind. The broader top-10 range is 58.0 points, so the table still separates the published systems.
12 models have been evaluated on OSWorld 2.0. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. OSWorld 2.0 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About OSWorld 2.0
Year
2026
Tasks
108 long-horizon computer-use workflows
Format
Interactive computer-use evaluation
Difficulty
Long-horizon professional workflows
OSWorld 2.0 expands computer-use evaluation to 108 long-horizon workflows that require state tracking, cross-source reasoning, visual-spatial precision, dynamic interaction, and verification. BenchLM stores the primary binary-completion score as a display-only agentic benchmark.
BenchLM freshness & provenance
Version
OSWorld 2.0 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does OSWorld 2.0 measure?
A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.
Which model scores highest on OSWorld 2.0?
GPT-5.6 Sol by OpenAI currently leads with a score of 62.6% on OSWorld 2.0.
How many models are evaluated on OSWorld 2.0?
12 AI models have been evaluated on OSWorld 2.0 on BenchLM.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.