Benchmark profile
MMLU-ProX
A multilingual extension of professional-level academic evaluation across many languages.
Data verifiedTop models on MMLU-ProX — July 23, 2026
As of July 23, 2026, Qwen3.7 Max leads the MMLU-ProX leaderboard with 87% , followed by Claude Opus 4.5 (85.7%) and Qwen3.7 Plus (85.4%).
Qwen3.7 Max
Alibaba
qwen3-7-max
Claude Opus 4.5
Anthropic
claude-opus-4-5
Qwen3.7 Plus
Alibaba
qwen3-7-plus
Leaderboard (12 models)
ScoreAccording to BenchLM.ai, Qwen3.7 Max leads the MMLU-ProX benchmark with a score of 87%, followed by Claude Opus 4.5 (85.7%) and Qwen3.7 Plus (85.4%). The top models are clustered within 1.6 points, suggesting this benchmark is nearing saturation for frontier models.
12 models have been evaluated on MMLU-ProX. The benchmark falls in the Multilingual category. This category carries a 7% weight in BenchLM.ai's overall scoring system. Within that category, MMLU-ProX contributes 100% of the category score, so strong performance here directly affects a model's overall ranking.
About MMLU-ProX
Year
2025
Tasks
Multilingual professional QA
Format
Multilingual multiple choice
Difficulty
Professional multilingual
MMLU-ProX expands multilingual evaluation beyond translated arithmetic, making it a better signal for broad cross-lingual reasoning and knowledge.
BenchLM freshness & provenance
Version
MMLU-ProX 2025
Refresh cadence
Static
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MMLU-ProX measure?
A multilingual extension of professional-level academic evaluation across many languages.
Which model scores highest on MMLU-ProX?
Qwen3.7 Max by Alibaba currently leads with a score of 87% on MMLU-ProX.
How many models are evaluated on MMLU-ProX?
12 AI models have been evaluated on MMLU-ProX on BenchLM.
Compare Top Models on MMLU-ProX
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.