Benchmark profile
MRCRv2
A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.
Data verifiedTop models on MRCRv2 — July 23, 2026
As of July 23, 2026, Sakana Fugu-Ultra leads the MRCRv2 leaderboard with 93.6% , followed by Qwen3.7 Plus (91.7%) and Qwen3.7 Max (90.4%).
Sakana Fugu-Ultra
Sakana AI
sakana-fugu-ultra
Qwen3.7 Plus
Alibaba
qwen3-7-plus
Qwen3.7 Max
Alibaba
qwen3-7-max
Leaderboard (7 models)
ScoreAccording to BenchLM.ai, Sakana Fugu-Ultra leads the MRCRv2 benchmark with a score of 93.6%, followed by Qwen3.7 Plus (91.7%) and Qwen3.7 Max (90.4%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.
7 models have been evaluated on MRCRv2. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. Within that category, MRCRv2 contributes 31% of the category score, so strong performance here directly affects a model's overall ranking.
About MRCRv2
Year
2025
Tasks
Long-context retrieval
Format
Multi-round long-context evaluation
Difficulty
Hard long-context
MRCRv2 is especially useful for models that compete on long context, since it checks whether they can retrieve the right information across long, multi-round interactions.
BenchLM freshness & provenance
Version
MRCRv2 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MRCRv2 measure?
A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.
Which model scores highest on MRCRv2?
Sakana Fugu-Ultra by Sakana AI currently leads with a score of 93.6% on MRCRv2.
How many models are evaluated on MRCRv2?
7 AI models have been evaluated on MRCRv2 on BenchLM.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.