Skip to main content

Benchmark profile

GeneBench-Pro

A multistage statistical-reasoning benchmark for genomics and biological-data analysis agents.

Data verified

Benchmark score on GeneBench-Pro — July 23, 2026

BenchLM mirrors the published score view for GeneBench-Pro. GPT-5.6 Sol leads the public snapshot at 28.7%. BenchLM does not use these results to rank models overall.

1 modelReasoningCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (1 model)

Score
1
GPT-5.6 SolOpenAI · Closed
28.7%

About GeneBench-Pro

Year

2026

Tasks

129 genomics statistical-analysis workflows

Format

Eval-level pass rate across dependent analysis decisions

Difficulty

Long-horizon scientific reasoning

GeneBench-Pro evaluates whether an agent chooses and executes the correct sequence of statistical-analysis decisions in genomics workflows. BenchLM stores the published result as display-only because it is a newly introduced, provider-developed benchmark with sparse public coverage.

BenchLM freshness & provenance

Version

GeneBench-Pro

Refresh cadence

Static

Staleness state

Current

Question availability

10 public tasks with provider-retained evaluation tasks

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does GeneBench-Pro measure?

A multistage statistical-reasoning benchmark for genomics and biological-data analysis agents.

Which model scores highest on GeneBench-Pro?

GPT-5.6 Sol by OpenAI currently leads with a score of 28.7% on GeneBench-Pro.

How many models are evaluated on GeneBench-Pro?

1 AI models have been evaluated on GeneBench-Pro on BenchLM.

Last updated: July 23, 2026 · BenchLM version GeneBench-Pro

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.