Skip to main content

Benchmark profile

Vals AIME (AIME)

Challenging national math exam given to top high-school students

Data verified

How BenchLM shows AIME

BenchLM mirrors the public Vals AI AIME leaderboard captured from https://www.vals.ai/benchmarks/aime and updated by Vals on April 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

AIME is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

96 Vals rows3 task viewspublic datasetTasks: Overall, AIME 2024, AIME 2025Display only

AIME score on AIME — April 16, 2026

BenchLM mirrors the published aime score view for AIME. Gemini 3.1 Pro Preview leads the public snapshot at 98.13% , followed by GPT-5.2 (96.88%) and Muse Spark (96.88%). BenchLM does not use these results to rank models overall.

96 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated April 16, 2026

AIME score table (96 models)

Score
2
GPT-5.2OpenAI
96.88%
3
96.88%
5
GPT-5.4OpenAI
96.67%
7
96.25%
9
95.63%
10
95.63%
11
95.63%
13
94.58%
14
GPT-5OpenAI
93.37%
15
GPT-5.1OpenAI
93.33%
16
GLM 4.7Zhipu AI
93.33%
17
GLM 4.6Zhipu AI
92.71%
18
GPT Oss 120bFireworks AI
92.60%
19
92.50%
20
92.29%
21
91.88%
22
GLM 5.1Zhipu AI
91.88%
23
91.67%
24
91.46%
25
91.25%
26
91.04%
27
Grok 4 0709SpaceXAI
90.56%
28
88.75%
29
88.75%
31
GLM 4.5Zhipu AI
86.67%
32
O3 MiniOpenAI
86.46%
33
86.04%
34
GPT Oss 20bFireworks AI
86.04%
36
Kimi K2 ThinkingMoonshot AI
85.42%
37
O3OpenAI
85.28%
39
84.58%
40
Qwen3 235b A22bFireworks AI
83.96%
41
O4 MiniOpenAI
83.67%
42
83.54%
45
81.18%
46
Qwen3 MaxAlibaba
81.04%
47
80.68%
49
77.92%
50
76.88%
52
DeepSeek R1Fireworks AI
73.96%
53
O1OpenAI
71.46%
55
DeepSeek V3p2Fireworks AI
64.79%
56
62.71%
57
60.69%
58
Grok 3SpaceXAI
58.75%
60
DeepSeek V3 0324Fireworks AI
52.20%
63
49.38%
65
44.24%
66
42.92%
67
42.29%
69
Claude Opus 4Anthropic
41.25%
70
GPT-4.1OpenAI
39.58%
71
38.54%
73
29.79%
74
DeepSeek V3Fireworks AI
27.50%
76
26.46%
78
25.21%
79
22.29%
81
18.75%
82
17.29%
84
Grok 2 1212SpaceXAI
15.21%
85
GPT-4oOpenAI
13.96%
86
13.33%
87
GPT-4oOpenAI
11.88%
88
11.46%
89
10.00%
91
9.17%
92
5.63%
93
3.54%
94
3.33%
95
0.42%
96
0.42%

The published AIME snapshot places Gemini 3.1 Pro Preview first at 98.13%. The third row is 1.25 points behind. The broader top-10 range is 2.50 points, so many of the published results sit in a relatively narrow band.

96 models have been evaluated on AIME. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. AIME is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AIME

Year

2026

Tasks

AIME math problems

Format

Accuracy score

Difficulty

Competition math

BenchLM mirrors the public Vals AI AIME leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

BenchLM freshness & provenance

Version

AIME 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AIME measure?

Challenging national math exam given to top high-school students

Which model leads the published AIME snapshot?

Gemini 3.1 Pro Preview currently leads the published AIME snapshot with 98.13% aime score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on AIME?

96 AI models are included in BenchLM's mirrored AIME snapshot, based on the public leaderboard captured on April 16, 2026.

Last updated: April 16, 2026 · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.