Benchmark profile
MMLU-Redux
A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.
Data verifiedBenchmark score on MMLU-Redux — July 23, 2026
BenchLM mirrors the published score view for MMLU-Redux. Claude Opus 4.5 leads the public snapshot at 96.6% , followed by Qwen3.7 Max (95%) and Qwen3.5 397B (94.9%). BenchLM does not use these results to rank models overall.
Claude Opus 4.5
Anthropic
claude-opus-4-5
Qwen3.7 Max
Alibaba
qwen3-7-max
Qwen3.5 397B
Alibaba
qwen3-5-397b
Benchmark score table (11 models)
ScoreThe published MMLU-Redux snapshot places Claude Opus 4.5 first at 96.6%. The third row is 1.7 points behind. The broader top-10 range is 18.5 points, so the table still separates the published systems.
11 models have been evaluated on MMLU-Redux. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. MMLU-Redux is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About MMLU-Redux
Year
2026
Tasks
Broad academic QA
Format
Multiple choice questions
Difficulty
Advanced general knowledge
MMLU-Redux is useful when MMLU itself has largely saturated. It acts as a broader knowledge sanity check with fresher or harder questions intended to preserve separation among strong general-purpose models.
BenchLM freshness & provenance
Version
MMLU-Redux 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MMLU-Redux measure?
A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.
Which model scores highest on MMLU-Redux?
Claude Opus 4.5 by Anthropic currently leads with a score of 96.6% on MMLU-Redux.
How many models are evaluated on MMLU-Redux?
11 AI models have been evaluated on MMLU-Redux on BenchLM.
Compare Top Models on MMLU-Redux
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.