Skip to main content

Benchmark profile

ResearchClawBench

An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.

Data verified

How BenchLM shows ResearchClawBench

BenchLM mirrors the official ResearchClawBench Pass@1 leaderboard snapshot. The source benchmark contains 40 tasks across 10 scientific domains and uses RADS average (Pass@1) as the primary metric.

ResearchClawBench gives agents related literature and raw data, hides the target paper, and grades how much of the scientific result they rediscover. The RADS scale treats 50 as matching the original paper and 70+ as surpassing it.

ResearchClawBench is display only on BenchLM. The rows combine a model, a research harness, execution budget, and long-horizon scientific workflow, so BenchLM does not use them as weighted base-model ranking inputs.

31 agent rows20 mapped ResearchHarness rows40 tasks10 domains6 Pass@5 rows trackedDisplay only

RADS average (Pass@1) on ResearchClawBench — 2026-06-29 snapshot

BenchLM mirrors the published rads average (pass@1) view for ResearchClawBench. Qiushi Engine leads the public snapshot at 30.2% , followed by Open Science (22.8%) and Claude Code (21.5%). BenchLM does not use these results to rank models overall.

31 modelsAgenticCurrentDisplay onlyUpdated 2026-06-29 snapshot

RADS average (Pass@1) table (31 models)

Score
1
Qiushi EngineInternScience
30.2%
2
Open ScienceInternScience
22.8%
3
Claude CodeInternScience
21.5%
4
Claude Opus 4.8Anthropic · Closed
21.1%
5
Claude Opus 4.7Anthropic · Closed
20.7%
6
GLM-5.2Z.AI · Open weight
20.7%
7
Claude Opus 4.6Anthropic · Closed
19.9%
8
MiniMax M3MiniMax · Open weight
19.8%
9
18.8%
10
Qwen3.7 MaxAlibaba · Closed
18.7%
11
Codex CLIInternScience
18.4%
12
GLM-5.1Z.AI · Open weight
18.2%
13
Kimi K2.6Moonshot AI · Open weight
18.0%
14
Qwen3.6 PlusAlibaba · Closed
18.0%
15
Gemini 3.5 FlashGoogle · Closed
17.9%
16
DeepSeek V4 ProDeepSeek · Open weight
17.1%
17
GPT-5.5OpenAI · Closed
17.0%
18
MiMo-V2.5Xiaomi · Closed
16.9%
19
OpenClawInternScience
16.6%
20
ResearchClawInternScience
16.3%
21
15.5%
22
MiMo-V2-ProXiaomi · Closed
15.3%
23
GPT-5.4OpenAI · Closed
15.3%
24
Qwen3.5 397BAlibaba · Open weight
14.2%
25
Kimi K2.5Moonshot AI · Open weight
14.0%
26
ARIS CodexInternScience
13.6%
27
Grok 4.1xAI · Closed
13.5%
28
Gemini 3.1 ProGoogle · Closed
13.3%
29
12.9%
30
NanobotInternScience
12.8%
31
Grok 4.3xAI · Closed
12.4%

The published ResearchClawBench snapshot places Qiushi Engine first at 30.2%. The third row is 8.7 points behind. The broader top-10 range is 11.5 points, so the table still separates the published systems.

31 models have been evaluated on ResearchClawBench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. ResearchClawBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ResearchClawBench

Year

2026

Tasks

40 tasks across 10 scientific domains

Format

End-to-end autonomous research evaluation with RADS scoring

Difficulty

Scientific research re-discovery

ResearchClawBench grades scientific agents with RADS, a rubric where 50 indicates matching the target paper and 70+ indicates surpassing it. BenchLM mirrors the official Pass@1 leaderboard as display-only because rows reflect a research-agent harness and long-horizon scientific workflow, not a normalized base-model-only comparison.

BenchLM freshness & provenance

Version

ResearchClawBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ResearchClawBench measure?

An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.

Which model leads the published ResearchClawBench snapshot?

Qiushi Engine currently leads the published ResearchClawBench snapshot with 30.2% rads average (pass@1). BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on ResearchClawBench?

31 AI models are included in BenchLM's mirrored ResearchClawBench snapshot, based on the public leaderboard captured on 2026-06-29 snapshot.

Last updated: 2026-06-29 snapshot · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.