Skip to main content

Benchmark profile

MRCRv2

A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.

Data verified

Top models on MRCRv2 — July 23, 2026

As of July 23, 2026, Sakana Fugu-Ultra leads the MRCRv2 leaderboard with 93.6% , followed by Qwen3.7 Plus (91.7%) and Qwen3.7 Max (90.4%).

7 modelsReasoning31% of category scoreCurrentUpdated July 23, 2026

Leaderboard (7 models)

Score
1
Sakana Fugu-UltraSakana AI · Closed
93.6%
2
Qwen3.7 PlusAlibaba · Closed
91.7%
3
Qwen3.7 MaxAlibaba · Closed
90.4%
4
Sakana FuguSakana AI · Closed
86.6%
5
Gemini 3.5 FlashGoogle · Closed
77.3%
6
Gemini 3.5 Flash-LiteGoogle · Closed
72.2%
7
Gemma 4 12BGoogle · Open weight
43.4%

According to BenchLM.ai, Sakana Fugu-Ultra leads the MRCRv2 benchmark with a score of 93.6%, followed by Qwen3.7 Plus (91.7%) and Qwen3.7 Max (90.4%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

7 models have been evaluated on MRCRv2. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. Within that category, MRCRv2 contributes 31% of the category score, so strong performance here directly affects a model's overall ranking.

About MRCRv2

Year

2025

Tasks

Long-context retrieval

Format

Multi-round long-context evaluation

Difficulty

Hard long-context

MRCRv2 is especially useful for models that compete on long context, since it checks whether they can retrieve the right information across long, multi-round interactions.

BenchLM freshness & provenance

Version

MRCRv2 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MRCRv2 measure?

A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.

Which model scores highest on MRCRv2?

Sakana Fugu-Ultra by Sakana AI currently leads with a score of 93.6% on MRCRv2.

How many models are evaluated on MRCRv2?

7 AI models have been evaluated on MRCRv2 on BenchLM.

Last updated: July 23, 2026 · BenchLM version MRCRv2 2025

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.