Skip to main content

BenchLM recommendation

Best AI Models for Web Research in 2026

Data verified

As of July 23, 2026, the top model in best ai models for web research on the BenchLM leaderboard is GPT-5.6 Sol with a score of 92.2.

Last verified: July 23, 2026

This reporting page isolates the web research slice of agentic performance. It prioritizes sourced benchmarks for browsing, evidence gathering, and multi-step web task completion rather than generic overall agent scores.

This page ranks models using only sourced web research benchmarks in the reporting family.

Bottom line: Web research agents need to browse, gather evidence, and synthesize findings. BrowseComp is the most predictive benchmark here.

GPT-5.6 Sol leads this ranking with a score of 92.2, followed by Kimi K3 (91.2) and GPT-5.5 Pro (90.1). The top three are separated by just a few points — any of them would perform well for this use case.

The best open-weight option is MiniMax M3 (ranked #11 with a score of 83.5). Proprietary models hold a clear advantage in this category, though open-weight options may suffice for less demanding use cases.

This ranking is based on provisional overall weighted scores across BenchLM.ai's scoring formula tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

How to choose

Full Rankings (32 models)

1
GPT-5.6 Sol
OpenAI·Proprietary·1M

92.2

sourced avg

2
Kimi K3
Moonshot AI·Pending·1.05M

91.2

sourced avg

3
GPT-5.5 Pro
OpenAI·Proprietary·1M

90.1

sourced avg

4
GPT-5.4 Pro
OpenAI·Proprietary·1.05M

89.3

sourced avg

5
Claude Mythos 5
Anthropic·Proprietary·1M+

88

sourced avg

6
GPT-5.6 Terra
OpenAI·Proprietary·1M

87.5

sourced avg

7
Claude Sonnet 5
Anthropic·Proprietary·1M

84.7

sourced avg

8
GPT-5.5
OpenAI·Proprietary·1M

84.4

sourced avg

9
Claude Opus 4.8
Anthropic·Proprietary·1M

84.3

sourced avg

10
Claude Opus 4.6
Anthropic·Proprietary·1M

83.7

sourced avg

11
MiniMax M3
MiniMax·Open Weight·1M

83.5

sourced avg

12
DeepSeek V4 Pro (Max)
DeepSeek·Open Weight·1M

83.4

sourced avg

13
GPT-5.6 Luna
OpenAI·Proprietary·1M

83.3

sourced avg

14
Kimi K2.6
Moonshot AI·Open Weight·256K

83.2

sourced avg

15
GPT-5.4
OpenAI·Proprietary·1.05M

82.7

sourced avg

16
DeepSeek V4 Pro (High)
DeepSeek·Open Weight·1M

80.4

sourced avg

17
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

79.3

sourced avg

18
Inkling
Thinking Machines Lab·Open Weight·1M

77.1

sourced avg

19
Step 3.7 Flash
StepFun·Open Weight·256K

75.8

sourced avg

20
Agents-A1
InternScience·Open Weight·262K

75.5

sourced avg

21
DeepSeek V4 Flash (Max)
DeepSeek·Open Weight·1M

73.2

sourced avg

22
GLM-5.1
Z.AI·Open Weight·203K

68

sourced avg

23
GPT-5.2
OpenAI·Proprietary·400K

65.8

sourced avg

24
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

63.8

sourced avg

25
Qwen3.5 397B
Alibaba·Open Weight·128K

62

sourced avg

26
Qwen3.5-27B
Alibaba·Open Weight·262K

61

sourced avg

27
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

61

sourced avg

28
Kimi K2.5
Moonshot AI·Open Weight·256K

60.6

sourced avg

29
Kimi K2.5 (Reasoning)
Moonshot AI·Proprietary·128K

60.6

sourced avg

30
DeepSeek V4 Flash (High)
DeepSeek·Open Weight·1M

53.5

sourced avg

31
GLM-4.7
Z.AI·Open Weight·200K

52

sourced avg

32
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

44.4

sourced avg

Key Takeaways

The top model on this sourced reporting-family slice is GPT-5.6 Sol by OpenAI with an average of 92.2.

The best open-weight model is MiniMax M3 at position #11.

32 models are listed with sourced benchmark coverage in this reporting family.

Score in Context

What these scores mean

This ranking averages sourced web-research benchmarks. It isolates the browsing and evidence-gathering slice of agentic performance.

Known limitations

Web research benchmarks test specific browsing patterns. Real-world web research also depends on access to search APIs, page rendering quality, and anti-bot measures that benchmarks do not capture.

Last updated: July 23, 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.