Skip to main content

Benchmark profile

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

Data verified

Top models on BrowseComp — July 23, 2026

As of July 23, 2026, GPT-5.6 Sol leads the BrowseComp leaderboard with 92.2% , followed by Kimi K3 (91.2%) and GPT-5.5 Pro (90.1%).

32 modelsAgentic28% of category scoreCurrentUpdated July 23, 2026

Leaderboard (32 models)

Score
1
GPT-5.6 SolOpenAI · Closed
92.2%
2
Kimi K3Moonshot AI · Closed
91.2%
3
GPT-5.5 ProOpenAI · Closed
90.1%
4
GPT-5.4 ProOpenAI · Closed
89.3%
5
Claude Mythos 5Anthropic · Closed
88%
6
GPT-5.6 TerraOpenAI · Closed
87.5%
7
Claude Sonnet 5Anthropic · Closed
84.7%
8
GPT-5.5OpenAI · Closed
84.4%
9
Claude Opus 4.8Anthropic · Closed
84.3%
10
Claude Opus 4.6Anthropic · Closed
83.7%
11
MiniMax M3MiniMax · Open weight
83.5%
12
DeepSeek V4 Pro (Max)DeepSeek · Open weight
83.4%
13
GPT-5.6 LunaOpenAI · Closed
83.3%
14
Kimi K2.6Moonshot AI · Open weight
83.2%
15
GPT-5.4OpenAI · Closed
82.7%
16
DeepSeek V4 Pro (High)DeepSeek · Open weight
80.4%
17
Claude Opus 4.7 (Adaptive)Anthropic · Closed
79.3%
18
InklingThinking Machines Lab · Open weight
77.1%
19
Step 3.7 FlashStepFun · Open weight
75.8%
20
Agents-A1InternScience · Open weight
75.5%
21
DeepSeek V4 Flash (Max)DeepSeek · Open weight
73.2%
22
GLM-5.1Z.AI · Open weight
68%
23
GPT-5.2OpenAI · Closed
65.8%
24
Qwen3.5-122B-A10BAlibaba · Open weight
63.8%
25
Qwen3.5 397BAlibaba · Open weight
62%
26
Qwen3.5-27BAlibaba · Open weight
61%
27
Qwen3.5-35B-A3BAlibaba · Open weight
61%
28
Kimi K2.5Moonshot AI · Open weight
60.6%
29
Kimi K2.5 (Reasoning)Moonshot AI · Closed
60.6%
30
DeepSeek V4 Flash (High)DeepSeek · Open weight
53.5%
31
GLM-4.7Z.AI · Open weight
52%
32
Nemotron 3 UltraNVIDIA · Open weight
44.4%

According to BenchLM.ai, GPT-5.6 Sol leads the BrowseComp benchmark with a score of 92.2%, followed by Kimi K3 (91.2%) and GPT-5.5 Pro (90.1%). The top models are clustered within 2.1 points, suggesting this benchmark is nearing saturation for frontier models.

32 models have been evaluated on BrowseComp. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Within that category, BrowseComp contributes 28% of the category score, so strong performance here directly affects a model's overall ranking.

About BrowseComp

Year

2025

Tasks

Research questions requiring browsing

Format

Web search and evidence synthesis

Difficulty

Hard web research

BrowseComp is designed to measure real web research behavior, not just latent world knowledge. It rewards models that can plan searches, inspect multiple pages, and avoid shallow answer synthesis.

BenchLM freshness & provenance

Version

BrowseComp 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does BrowseComp measure?

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

Which model scores highest on BrowseComp?

GPT-5.6 Sol by OpenAI currently leads with a score of 92.2% on BrowseComp.

How many models are evaluated on BrowseComp?

32 AI models have been evaluated on BrowseComp on BenchLM.

Last updated: July 23, 2026 · BenchLM version BrowseComp 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.