Skip to main content

Benchmark profile

Vals-hosted MMLU-Pro mirror (Vals MMLU-Pro mirror)

Vals AI hosted MMLU-Pro view with subject-level task splits.

Data verified

How BenchLM shows Vals MMLU-Pro mirror

BenchLM mirrors the public Vals AI Vals MMLU-Pro mirror leaderboard captured from https://www.vals.ai/benchmarks/mmlu_pro and updated by Vals on July 19, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals MMLU-Pro mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

122 Vals rows15 task viewspublic datasetTasks: Overall, Biology, Business, Chemistry, Computer ScienceDisplay only

Vals MMLU-Pro mirror score on Vals MMLU-Pro mirror — July 19, 2026

BenchLM mirrors the published vals mmlu-pro mirror score view for Vals MMLU-Pro mirror. Claude Fable 5 leads the public snapshot at 91.50% , followed by Gemini 3.1 Pro Preview (90.99%) and Gemini 3 Pro Preview (90.10%). BenchLM does not use these results to rank models overall.

122 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated July 19, 2026

Vals MMLU-Pro mirror score table (122 models)

Score
1
91.50%
4
89.87%
5
89.58%
6
89.52%
7
89.31%
8
Grok 4.5SpaceXAI
89.22%
9
89.11%
10
89.10%
11
88.73%
13
GPT-5.5OpenAI
88.14%
14
Kimi K3Moonshot AI
87.97%
16
87.67%
17
Kimi K2.6Moonshot AI
87.57%
18
87.55%
19
GPT-5.4OpenAI
87.48%
21
87.34%
22
87.32%
24
87.25%
25
87.21%
26
87.18%
27
87.05%
28
GLM 5.1Zhipu AI
86.90%
29
GLM 5.2Zhipu AI
86.71%
30
86.66%
31
GPT-5OpenAI
86.54%
32
GPT-5.1OpenAI
86.38%
33
InklingThinkingmachines
86.30%
34
86.25%
36
GPT-5.2OpenAI
86.23%
37
Claude Opus 4Anthropic
86.17%
38
86.04%
39
86.03%
40
85.91%
41
Grok 4.3SpaceXAI
85.84%
43
O3OpenAI
85.59%
44
85.59%
45
Grok 4 0709SpaceXAI
85.30%
46
Qwen3 MaxAlibaba
84.98%
47
84.92%
48
84.59%
49
84.55%
50
Qwen3 MaxAlibaba
84.36%
51
MiniMax M3MiniMax
84.22%
52
84.18%
53
84.06%
58
83.54%
59
O1OpenAI
83.49%
60
DeepSeek R1Fireworks AI
83.18%
61
DeepSeek V3p2Fireworks AI
83.06%
62
Mimo V2.5Xiaomi
82.93%
63
GLM 4.7Zhipu AI
82.74%
65
82.23%
66
GLM 4.6Zhipu AI
82.20%
68
Qwen3 235b A22bFireworks AI
81.25%
69
GLM 4.5Zhipu AI
81.22%
70
Kimi K2 ThinkingMoonshot AI
81.07%
71
80.66%
72
O4 MiniOpenAI
80.56%
73
GPT-4.1OpenAI
80.50%
74
80.43%
75
80.09%
77
Grok 3SpaceXAI
79.95%
78
79.82%
79
79.70%
80
DeepSeek V3 0324Fireworks AI
79.47%
81
79.43%
82
79.42%
83
79.39%
84
GPT Oss 120bFireworks AI
79.17%
87
O3 MiniOpenAI
78.69%
89
78.40%
90
77.38%
91
77.22%
92
77.17%
93
76.07%
94
Grok 2 1212SpaceXAI
75.47%
95
75.33%
96
75.29%
97
75.29%
99
GPT-4oOpenAI
74.13%
100
DeepSeek V3Fireworks AI
73.82%
101
GPT-4oOpenAI
72.56%
102
GPT Oss 20bFireworks AI
71.64%
104
70.34%
106
69.71%
109
Laguna Xs.2Poolside
69.41%
110
69.17%
111
68.66%
112
66.02%
113
65.61%
114
64.44%
115
64.12%
116
63.48%
117
Laguna M.1Poolside
63.48%
118
62.73%
119
62.13%
120
49.78%
121
44.00%
122
30.28%

The published Vals MMLU-Pro mirror snapshot places Claude Fable 5 first at 91.50%. The third row is 1.40 points behind. The broader top-10 range is 2.40 points, so many of the published results sit in a relatively narrow band.

122 models have been evaluated on Vals MMLU-Pro mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals MMLU-Pro mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals MMLU-Pro mirror

Year

2026

Tasks

MMLU-Pro subject splits

Format

Accuracy score

Difficulty

Professional academic reasoning

BenchLM keeps this Vals-hosted MMLU-Pro table separate from canonical MMLU-Pro source records.

BenchLM freshness & provenance

Version

Vals MMLU-Pro mirror 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Vals MMLU-Pro mirror measure?

Vals AI hosted MMLU-Pro view with subject-level task splits.

Which model leads the published Vals MMLU-Pro mirror snapshot?

Claude Fable 5 currently leads the published Vals MMLU-Pro mirror snapshot with 91.50% vals mmlu-pro mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals MMLU-Pro mirror?

122 AI models are included in BenchLM's mirrored Vals MMLU-Pro mirror snapshot, based on the public leaderboard captured on July 19, 2026.

Last updated: July 19, 2026 · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.