Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

Data verified 24 confirmed releases in the last 30 daysStart free brief

Top models on BrowseComp — August 22, 2026

As of August 22, 2026, GPT-5.6 Sol leads the BrowseComp leaderboard with 92.2% , followed by Kimi K3 (91.2%) and Claude Opus 5 (90.8%).

40 modelsAgentic28% of category scoreCurrentUpdated August 22, 2026

Leaderboard (40 models)

Score
1
GPT-5.6 SolOpenAI · Closed
92.2%
2
Kimi K3Moonshot AI · Closed
91.2%
3
Claude Opus 5Anthropic · Closed
90.8%
4
GPT-5.5 ProOpenAI · Closed
90.1%
5
GPT-5.4 ProOpenAI · Closed
89.3%
6
Claude Mythos 5Anthropic · Closed
88%
7
GPT-5.6 TerraOpenAI · Closed
87.5%
8
Ornith-1.5-397BOrnith AI · Open weight
86.6%
9
Claude Sonnet 5Anthropic · Closed
84.7%
10
GPT-5.5OpenAI · Closed
84.4%
11
Claude Opus 4.8Anthropic · Closed
84.3%
12
Claude Opus 4.6Anthropic · Closed
83.7%
13
MiniMax M3MiniMax · Open weight
83.5%
14
DeepSeek V4 Pro 0813DeepSeek · Closed
83.4%
15
GPT-5.6 LunaOpenAI · Closed
83.3%
16
dots3-note PreviewDots Studio · Open weight
83.3%
17
Kimi K2.6Moonshot AI · Open weight
83.2%
18
GPT-5.4OpenAI · Closed
82.7%
19
DeepSeek V4 Pro (High)DeepSeek · Open weight
80.4%
20
Claude Opus 4.7 (Adaptive)Anthropic · Closed
79.3%
21
Inkling-SmallThinking Machines Lab · Open weight
77.4%
22
InklingThinking Machines Lab · Open weight
77.1%
23
Step 3.7 FlashStepFun · Open weight
75.8%
24
Agents-A1InternScience · Open weight
75.5%
25
DeepSeek V4 Flash 0731DeepSeek · Closed
73.2%
26
Ling 3.0 FlashInclusionAI · Open weight
72.2%
27
GLM-5.1Z.AI · Open weight
68%
28
Ornith-1.5-35B-A3BOrnith AI · Open weight
67.6%
29
GPT-5.2OpenAI · Closed
65.8%
30
Qwen3.5-122B-A10BAlibaba · Open weight
63.8%
31
Qwen3.5 397BAlibaba · Open weight
62%
32
Qwen3.5-27BAlibaba · Open weight
61%
33
Qwen3.5-35B-A3BAlibaba · Open weight
61%
34
Kimi K2.5Moonshot AI · Open weight
60.6%
35
Kimi K2.5 (Reasoning)Moonshot AI · Closed
60.6%
36
Ornith-1.5-9BOrnith AI · Open weight
56.4%
37
DeepSeek V4 Flash (High)DeepSeek · Closed
53.5%
38
GLM-4.7Z.AI · Open weight
52%
39
Nemotron 3 UltraNVIDIA · Open weight
44.4%
40
36.8%

According to BenchLM.ai, GPT-5.6 Sol leads the BrowseComp benchmark with a score of 92.2%, followed by Kimi K3 (91.2%) and Claude Opus 5 (90.8%). The top models are clustered within 1.4 points, suggesting this benchmark is nearing saturation for frontier models.

40 models have been evaluated on BrowseComp. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Within that category, BrowseComp contributes 28% of the category score, so strong performance here directly affects a model's overall ranking.

About BrowseComp

Year

2025

Tasks

Research questions requiring browsing

Format

Web search and evidence synthesis

Difficulty

Hard web research

BrowseComp is designed to measure real web research behavior, not just latent world knowledge. It rewards models that can plan searches, inspect multiple pages, and avoid shallow answer synthesis.

BenchLM freshness & provenance

Version

BrowseComp 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does BrowseComp measure?

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

Which model scores highest on BrowseComp?

GPT-5.6 Sol by OpenAI currently leads with a score of 92.2% on BrowseComp.

How many models are evaluated on BrowseComp?

40 AI models have been evaluated on BrowseComp on BenchLM.

Last updated: August 22, 2026 · BenchLM version BrowseComp 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.