Skip to main content

Benchmark profile

WebArena-Verified Browser Agent Benchmark (WebArena-Verified)

WebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task descriptions, reference answers, and evaluators, and replaces nondeterministic judging with deterministic checks where possible.

Data verified

How to read this benchmark

Editorial review by Glevd · 2026-07-15

Treat each score as a result for the complete model-agent system, not the base model alone. Compare rows only when the full suite or Hard subset, environment revision, scaffold, prompts, DOM or screenshot interface, action and tool budget, evaluator, and pass@1 or repeat policy match.

Operator receipt: 0 sourced rows are currently displayable on this page.

Honest limit: The live ledger currently contains one exact provider-reported row, so it cannot establish a broad market leader. That run used the full 812-task suite, OpenClaw, hybrid DOM and screenshot interaction, a 150-tool-call cap, and mean pass@1. The fixed environments still omit much of today's open-web drift, production authentication, anti-bot behavior, permission boundaries, latency, safety controls, and cost.

About WebArena-Verified

Year

2025

Tasks

812 verified tasks; separate 258-task Hard subset

Format

Deterministic end-state task success

Difficulty

Audited stateful browser work

The full release contains 812 verified tasks across six self-hosted web environments. A separate 258-task Hard subset supports lower-cost evaluation. The maintainers manually reviewed every task, reference answer, and evaluator, and removed LLM-as-a-judge and substring-matching checks in favor of deterministic scoring.

BenchLM freshness & provenance

Version

WebArena-Verified 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What changed in WebArena-Verified?

The maintainers manually reviewed and corrected every task, reference answer, and evaluator. They also replaced LLM judging and substring matching with deterministic, type-aware checks where possible. The full release has 812 tasks, with a separate 258-task Hard subset.

Are WebArena-Verified scores directly comparable?

Only when the dataset subset, environment revision, agent scaffold, prompts, observation and action interfaces, tool or step budget, evaluator, and attempt policy match. A score belongs to that complete evaluation setup, not the base model alone.

Does WebArena-Verified replace live workflow testing?

No. Its self-hosted environments support repeatable evaluation, but they do not reproduce every current website, login flow, anti-bot system, permission boundary, latency constraint, safety control, or operating cost.

Last updated: July 23, 2026 · BenchLM version WebArena-Verified 2025

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.