Benchmark profile
ScreenSpot Pro
A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.
Data verifiedThe public ScreenSpot Pro snapshot ranks Claude Opus 4.8 first at 87.9%, ahead of GPT-5.4 (85.4%) and Gemini 3.1 Pro (84.4%) among 15 tested models. We mirror the table as display-only evidence; it does not affect overall rankings.
How to read this leaderboard
Editorial review by Glevd · 2026-07-15
Use ScreenSpot Pro to judge published GUI-grounding systems after checking how each system handles the image. Direct coordinate prediction, cropping, visual search, a planner, or Python tools can change the result without changing the underlying base model.
Operator receipt: 15 sourced rows are currently displayable on this page; the leading published row is Claude Opus 4.8 at 87.9%.
Honest limit: ScreenSpot Pro stops at localization on a static screenshot. It does not test clicking, typing, state tracking, recovery, or completion of a multi-step workflow, and published rows with different tool or search setups are not a clean model-only comparison.
Benchmark score on ScreenSpot Pro — July 23, 2026
BenchLM mirrors the published score view for ScreenSpot Pro. Claude Opus 4.8 leads the public snapshot at 87.9% , followed by GPT-5.4 (85.4%) and Gemini 3.1 Pro (84.4%). BenchLM does not use these results to rank models overall.
Claude Opus 4.8
Anthropic
claude-opus-4-8
GPT-5.4
OpenAI
gpt-5-4
Gemini 3.1 Pro
gemini-3-1-pro
Benchmark score table (15 models)
ScoreThe published ScreenSpot Pro snapshot places Claude Opus 4.8 first at 87.9%. The third row is 3.5 points behind. The broader top-10 range is 21.8 points, so the table still separates the published systems.
15 models have been evaluated on ScreenSpot Pro. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. ScreenSpot Pro is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About ScreenSpot Pro
Year
2025
Tasks
1,581 grounding instructions
Format
Static interface element localization
Difficulty
Professional GUI grounding
The benchmark spans 23 applications, six application categories, and three operating systems. It scores whether a system localizes the requested text or icon in a static professional screenshot, using micro-average accuracy on the official leaderboard.
BenchLM freshness & provenance
Version
ScreenSpot Pro 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does ScreenSpot Pro measure?
ScreenSpot Pro measures whether a model can map a natural-language instruction to the correct target in a full-screen, high-resolution professional interface. Its 1,581 examples span 23 applications, six application categories, and three operating systems. The score is grounding accuracy, not end-to-end task completion.
What does a high ScreenSpot Pro score mean?
A high score means the published system localized more requested interface elements under its reported setup. It does not prove the same base model will click correctly inside an agent. Cropping, visual search, planners, Python tools, image resolution, and decoding policy can materially change the result.
Can ScreenSpot Pro pick the best computer-use agent?
No. ScreenSpot Pro tests static-screen grounding, which is one prerequisite for computer use. It does not test state tracking, typing, recovery after a bad action, or completion of a multi-step workflow. Pair it with OSWorld Verified and an application-specific operator test before choosing an agent.
Compare Top Models on ScreenSpot Pro
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.