Benchmark profile
WebArena Web Agent Benchmark (WebArena)
WebArena tests whether a browser-agent system can complete 812 long-horizon tasks inside self-hosted replicas of functional websites. It checks the requested end state, so a result reflects the model, agent scaffold, browser interface, action budget, and evaluator together—not the base model alone.
Data verifiedHow to read this benchmark
Editorial review by Glevd · 2026-07-15
Read this page as the original benchmark-method guide. Compare original WebArena rows only when the environment revision, model, agent scaffold, prompts, observation and action interfaces, step or tool-call budget, evaluator, and pass@1 or repeat policy match. Do not mix WebArena-Verified results into this lane.
Operator receipt: 0 sourced rows are currently displayable on this page.
Honest limit: No original-WebArena row currently has an exact-source attachment that BenchLM can display. The route therefore explains the method without publishing the 15 raw manual entries as a leaderboard. Fixed self-hosted sites do not directly test today's open web, layout drift, anti-bot systems, production login flows, real-world identity and permission boundaries, safety controls, latency, or cost.
About WebArena
Year
2024
Tasks
812 long-horizon browser tasks
Format
End-state task success
Difficulty
Stateful multi-site browser work
Tasks start from high-level natural-language goals across shopping, forums, collaborative software development, content management, maps, and knowledge sites. Evaluation uses answer matching or programmatic checks of site state, allowing more than one valid action path. This route owns original WebArena methodology. The separately audited WebArena-Verified release has its own result route.
BenchLM freshness & provenance
Version
WebArena 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does WebArena measure?
WebArena measures whether a browser-agent system can complete 812 long-horizon tasks in self-hosted website replicas. Tasks cover shopping, forums, software collaboration, content management, maps, and knowledge work. Success is checked against the requested answer or resulting site state, so multiple valid action paths can pass.
Are WebArena scores directly comparable?
No. Match original WebArena or WebArena-Verified, full suite or subset, environment revision, agent scaffold, prompts, observation and action interfaces, tool or step budget, evaluator, and attempt policy. A model tested with different browser tools or repeated attempts is not part of the same controlled comparison.
Can WebArena choose the best browser agent?
Not by itself. WebArena is a useful signal for stateful browser work, but its fixed self-hosted sites do not reproduce today's changing web, production login flows, anti-bot systems, real-world identity and permission boundaries, latency, cost, or safety controls. Use matched WebArena results alongside live workflow trials and broader agentic evidence.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.