Skip to main content

Benchmark profile

Massive Multi-discipline Multimodal Understanding Pro (MMMU-Pro)

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

Data verified

Top models on MMMU-Pro — July 23, 2026

As of July 23, 2026, GPT-5.4 Pro leads the MMMU-Pro leaderboard with 94% , followed by Gemini 3.1 Pro (83.9%) and Gemini 3.5 Flash (83.6%).

34 modelsMultimodal & Grounded45% of category scoreRefreshingUpdated July 23, 2026

Leaderboard (34 models)

Score
1
GPT-5.4 ProOpenAI · Closed
94%
2
Gemini 3.1 ProGoogle · Closed
83.9%
3
Gemini 3.5 FlashGoogle · Closed
83.6%
4
GPT-5.6 SolOpenAI · Closed
83%
5
Kimi K3Moonshot AI · Closed
81.6%
6
GPT-5.5OpenAI · Closed
81.2%
7
GPT-5.4OpenAI · Closed
81.2%
8
Gemini 3 ProGoogle · Closed
81%
9
GPT-5.6 TerraOpenAI · Closed
80.7%
10
Muse SparkMeta · Closed
80.4%
11
GPT-5.2OpenAI · Closed
79.5%
12
Kimi K2.6Moonshot AI · Open weight
79.4%
13
Qwen3.7 PlusAlibaba · Closed
79%
14
Qwen3.5 397BAlibaba · Open weight
79%
15
Qwen3.6 PlusAlibaba · Closed
78.8%
16
Kimi K2.5Moonshot AI · Open weight
78.5%
17
Kimi K2.5 (Reasoning)Moonshot AI · Closed
78.5%
18
GPT-5.6 LunaOpenAI · Closed
78.4%
19
MiniMax M3MiniMax · Open weight
78.1%
20
Grok 4.3xAI · Closed
78.1%
21
MiMo-V2.5Xiaomi · Closed
77.9%
22
Claude Opus 4.6Anthropic · Closed
77.3%
23
Gemma 4 31BGoogle · Open weight
76.9%
24
GPT-5.4 miniOpenAI · Closed
76.6%
25
Qwen3.6-27BAlibaba · Open weight
75.8%
26
Qwen3.6-35B-A3BAlibaba · Open weight
75.3%
27
Grok 4.20xAI · Closed
75.2%
28
Gemma 4 26B A4BGoogle · Open weight
73.8%
29
InklingThinking Machines Lab · Open weight
73.5%
30
Interfaze BetaInterfaze · Closed
71.1%
31
Claude Opus 4.5Anthropic · Closed
70.6%
32
Gemma 4 12BGoogle · Open weight
69.1%
33
GPT-5.4 nanoOpenAI · Closed
66.1%
34
Command A+Cohere · Open weight
63%

According to BenchLM.ai, GPT-5.4 Pro leads the MMMU-Pro benchmark with a score of 94%, followed by Gemini 3.1 Pro (83.9%) and Gemini 3.5 Flash (83.6%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.

34 models have been evaluated on MMMU-Pro. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, MMMU-Pro contributes 45% of the category score, so strong performance here directly affects a model's overall ranking.

About MMMU-Pro

Year

2024

Tasks

Multimodal academic reasoning

Format

Image + text question answering

Difficulty

Frontier multimodal

MMMU-Pro extends the original MMMU setup with more difficult multimodal questions and stronger separation at the top end of the model market.

BenchLM freshness & provenance

Version

MMMU-Pro 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MMMU-Pro measure?

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

Which model scores highest on MMMU-Pro?

GPT-5.4 Pro by OpenAI currently leads with a score of 94% on MMMU-Pro.

How many models are evaluated on MMMU-Pro?

34 AI models have been evaluated on MMMU-Pro on BenchLM.

Last updated: July 23, 2026 · BenchLM version MMMU-Pro 2024

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.