Skip to main content

Benchmark profile

WebArena Web Agent Benchmark (WebArena)

WebArena tests whether a browser-agent system can complete 812 long-horizon tasks inside self-hosted replicas of functional websites. It checks the requested end state, so a result reflects the model, agent scaffold, browser interface, action budget, and evaluator together—not the base model alone.

Data verified

How to read this benchmark

Editorial review by Glevd · 2026-07-15

Read this page as the original benchmark-method guide. Compare original WebArena rows only when the environment revision, model, agent scaffold, prompts, observation and action interfaces, step or tool-call budget, evaluator, and pass@1 or repeat policy match. Do not mix WebArena-Verified results into this lane.

Operator receipt: 0 sourced rows are currently displayable on this page.

Honest limit: No original-WebArena row currently has an exact-source attachment that BenchLM can display. The route therefore explains the method without publishing the 15 raw manual entries as a leaderboard. Fixed self-hosted sites do not directly test today's open web, layout drift, anti-bot systems, production login flows, real-world identity and permission boundaries, safety controls, latency, or cost.

About WebArena

Year

2024

Tasks

812 long-horizon browser tasks

Format

End-state task success

Difficulty

Stateful multi-site browser work

Tasks start from high-level natural-language goals across shopping, forums, collaborative software development, content management, maps, and knowledge sites. Evaluation uses answer matching or programmatic checks of site state, allowing more than one valid action path. This route owns original WebArena methodology. The separately audited WebArena-Verified release has its own result route.

BenchLM freshness & provenance

Version

WebArena 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does WebArena measure?

WebArena measures whether a browser-agent system can complete 812 long-horizon tasks in self-hosted website replicas. Tasks cover shopping, forums, software collaboration, content management, maps, and knowledge work. Success is checked against the requested answer or resulting site state, so multiple valid action paths can pass.

Are WebArena scores directly comparable?

No. Match original WebArena or WebArena-Verified, full suite or subset, environment revision, agent scaffold, prompts, observation and action interfaces, tool or step budget, evaluator, and attempt policy. A model tested with different browser tools or repeated attempts is not part of the same controlled comparison.

Can WebArena choose the best browser agent?

Not by itself. WebArena is a useful signal for stateful browser work, but its fixed self-hosted sites do not reproduce today's changing web, production login flows, anti-bot systems, real-world identity and permission boundaries, latency, cost, or safety controls. Use matched WebArena results alongside live workflow trials and broader agentic evidence.

Last updated: July 23, 2026 · BenchLM version WebArena 2024

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.