Benchmark profile
WebArena-Verified Browser Agent Benchmark (WebArena-Verified)
WebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task descriptions, reference answers, and evaluators, and replaces nondeterministic judging with deterministic checks where possible.
Data verifiedHow to read this benchmark
Editorial review by Glevd · 2026-07-15
Treat each score as a result for the complete model-agent system, not the base model alone. Compare rows only when the full suite or Hard subset, environment revision, scaffold, prompts, DOM or screenshot interface, action and tool budget, evaluator, and pass@1 or repeat policy match.
Operator receipt: 0 sourced rows are currently displayable on this page.
Honest limit: The live ledger currently contains one exact provider-reported row, so it cannot establish a broad market leader. That run used the full 812-task suite, OpenClaw, hybrid DOM and screenshot interaction, a 150-tool-call cap, and mean pass@1. The fixed environments still omit much of today's open-web drift, production authentication, anti-bot behavior, permission boundaries, latency, safety controls, and cost.
About WebArena-Verified
Year
2025
Tasks
812 verified tasks; separate 258-task Hard subset
Format
Deterministic end-state task success
Difficulty
Audited stateful browser work
The full release contains 812 verified tasks across six self-hosted web environments. A separate 258-task Hard subset supports lower-cost evaluation. The maintainers manually reviewed every task, reference answer, and evaluator, and removed LLM-as-a-judge and substring-matching checks in favor of deterministic scoring.
BenchLM freshness & provenance
Version
WebArena-Verified 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What changed in WebArena-Verified?
The maintainers manually reviewed and corrected every task, reference answer, and evaluator. They also replaced LLM judging and substring matching with deterministic, type-aware checks where possible. The full release has 812 tasks, with a separate 258-task Hard subset.
Are WebArena-Verified scores directly comparable?
Only when the dataset subset, environment revision, agent scaffold, prompts, observation and action interfaces, tool or step budget, evaluator, and attempt policy match. A score belongs to that complete evaluation setup, not the base model alone.
Does WebArena-Verified replace live workflow testing?
No. Its self-hosted environments support repeatable evaluation, but they do not reproduce every current website, login flow, anti-bot system, permission boundary, latency constraint, safety control, or operating cost.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.