Benchmark profile
ResearchClawBench
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
Data verifiedHow BenchLM shows ResearchClawBench
BenchLM mirrors the official ResearchClawBench Pass@1 leaderboard snapshot. The source benchmark contains 40 tasks across 10 scientific domains and uses RADS average (Pass@1) as the primary metric.
ResearchClawBench gives agents related literature and raw data, hides the target paper, and grades how much of the scientific result they rediscover. The RADS scale treats 50 as matching the original paper and 70+ as surpassing it.
ResearchClawBench is display only on BenchLM. The rows combine a model, a research harness, execution budget, and long-horizon scientific workflow, so BenchLM does not use them as weighted base-model ranking inputs.
RADS average (Pass@1) on ResearchClawBench — 2026-06-29 snapshot
BenchLM mirrors the published rads average (pass@1) view for ResearchClawBench. Qiushi Engine leads the public snapshot at 30.2% , followed by Open Science (22.8%) and Claude Code (21.5%). BenchLM does not use these results to rank models overall.
Qiushi Engine
InternScience
gpt-5.5
Open Science
InternScience
claude-opus-4.8
Claude Code
InternScience
claude-opus-4-6
RADS average (Pass@1) table (31 models)
ScoreThe published ResearchClawBench snapshot places Qiushi Engine first at 30.2%. The third row is 8.7 points behind. The broader top-10 range is 11.5 points, so the table still separates the published systems.
31 models have been evaluated on ResearchClawBench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. ResearchClawBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About ResearchClawBench
Year
2026
Tasks
40 tasks across 10 scientific domains
Format
End-to-end autonomous research evaluation with RADS scoring
Difficulty
Scientific research re-discovery
ResearchClawBench grades scientific agents with RADS, a rubric where 50 indicates matching the target paper and 70+ indicates surpassing it. BenchLM mirrors the official Pass@1 leaderboard as display-only because rows reflect a research-agent harness and long-horizon scientific workflow, not a normalized base-model-only comparison.
BenchLM freshness & provenance
Version
ResearchClawBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does ResearchClawBench measure?
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
Which model leads the published ResearchClawBench snapshot?
Qiushi Engine currently leads the published ResearchClawBench snapshot with 30.2% rads average (pass@1). BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on ResearchClawBench?
31 AI models are included in BenchLM's mirrored ResearchClawBench snapshot, based on the public leaderboard captured on 2026-06-29 snapshot.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.