Skip to main content

Benchmark profile

MMLU-ProX

A multilingual extension of professional-level academic evaluation across many languages.

Data verified

Top models on MMLU-ProX — July 23, 2026

As of July 23, 2026, Qwen3.7 Max leads the MMLU-ProX leaderboard with 87% , followed by Claude Opus 4.5 (85.7%) and Qwen3.7 Plus (85.4%).

12 modelsMultilingual100% of category scoreCurrentUpdated July 23, 2026

Leaderboard (12 models)

Score
1
Qwen3.7 MaxAlibaba · Closed
87%
2
Claude Opus 4.5Anthropic · Closed
85.7%
3
Qwen3.7 PlusAlibaba · Closed
85.4%
4
Qwen3.6 PlusAlibaba · Closed
84.7%
5
Qwen3.5 397BAlibaba · Open weight
84.7%
6
GLM-5Z.AI · Open weight
83.1%
7
Nemotron 3 UltraNVIDIA · Open weight
83%
8
Kimi K2.5Moonshot AI · Open weight
82.3%
9
Qwen3.5-122B-A10BAlibaba · Open weight
82.2%
10
Qwen3.5-27BAlibaba · Open weight
82.2%
11
Qwen3.5-35B-A3BAlibaba · Open weight
81%
12
Qwen3 235B 2507Alibaba · Open weight
79.4%

According to BenchLM.ai, Qwen3.7 Max leads the MMLU-ProX benchmark with a score of 87%, followed by Claude Opus 4.5 (85.7%) and Qwen3.7 Plus (85.4%). The top models are clustered within 1.6 points, suggesting this benchmark is nearing saturation for frontier models.

12 models have been evaluated on MMLU-ProX. The benchmark falls in the Multilingual category. This category carries a 7% weight in BenchLM.ai's overall scoring system. Within that category, MMLU-ProX contributes 100% of the category score, so strong performance here directly affects a model's overall ranking.

About MMLU-ProX

Year

2025

Tasks

Multilingual professional QA

Format

Multilingual multiple choice

Difficulty

Professional multilingual

MMLU-ProX expands multilingual evaluation beyond translated arithmetic, making it a better signal for broad cross-lingual reasoning and knowledge.

BenchLM freshness & provenance

Version

MMLU-ProX 2025

Refresh cadence

Static

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MMLU-ProX measure?

A multilingual extension of professional-level academic evaluation across many languages.

Which model scores highest on MMLU-ProX?

Qwen3.7 Max by Alibaba currently leads with a score of 87% on MMLU-ProX.

How many models are evaluated on MMLU-ProX?

12 AI models have been evaluated on MMLU-ProX on BenchLM.

Last updated: July 23, 2026 · BenchLM version MMLU-ProX 2025

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.