Skip to main content

Benchmark profile

FrontierMath v2 Tiers 1-3 (FrontierMath v2 (Tiers 1-3))

Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.

Data verified

Top models on FrontierMath v2 (Tiers 1-3) — July 23, 2026

As of July 23, 2026, GPT-5.6 Sol leads the FrontierMath v2 (Tiers 1-3) leaderboard with 89.000% , followed by GPT-5.6 Terra (84.900%) and GPT-5.6 Luna (78.600%).

53 modelsMath30% of category scoreCurrentUpdated July 23, 2026

Leaderboard (53 models)

Score
1
GPT-5.6 SolOpenAI · Closed
89.000%
2
GPT-5.6 TerraOpenAI · Closed
84.900%
3
GPT-5.6 LunaOpenAI · Closed
78.600%
4
GPT-5.5OpenAI · Closed
51.700%
5
GPT-5.5 ProOpenAI · Closed
51.000%
6
GPT-5.4 ProOpenAI · Closed
50.000%
7
GPT-5.4OpenAI · Closed
47.600%
8
Claude Opus 4.8Anthropic · Closed
47.241%
9
Claude Opus 4.7Anthropic · Closed
43.793%
10
Claude Opus 4.6Anthropic · Closed
40.700%
11
GPT-5.2OpenAI · Closed
40.700%
12
Muse SparkMeta · Closed
39.000%
13
Gemini 3.5 FlashGoogle · Closed
38.966%
14
Kimi K2.6Moonshot AI · Open weight
38.966%
15
Gemini 3 ProGoogle · Closed
37.600%
16
Gemini 3.1 ProGoogle · Closed
36.900%
17
Gemini 3 FlashGoogle · Closed
35.640%
18
GLM-5.1Z.AI · Open weight
33.448%
19
Claude Sonnet 4.6Anthropic · Closed
32.400%
20
GPT-5.1OpenAI · Closed
31.034%
21
GPT-5.4 miniOpenAI · Closed
28.280%
22
Kimi K2.5Moonshot AI · Open weight
27.900%
23
GPT-5 miniOpenAI · Closed
27.241%
24
Qwen3.6 PlusAlibaba · Closed
26.207%
25
GPT-5.4 nanoOpenAI · Closed
25.860%
26
o4-mini (high)OpenAI · Closed
24.828%
27
Qwen 3.6 Max (preview)Alibaba · Closed
23.103%
28
DeepSeek V3.2DeepSeek · Open weight
22.100%
29
Kimi K2Moonshot AI · Closed
21.404%
30
Qwen3.5 PlusAlibaba · Closed
21.034%
31
Claude Opus 4.5Anthropic · Closed
20.690%
32
Grok 4xAI · Closed
19.655%
33
o3OpenAI · Closed
18.685%
34
GLM-5Z.AI · Open weight
16.434%
35
Gemini 2.5 ProGoogle · Closed
14.138%
36
Claude Sonnet 4.5Anthropic · Closed
13.495%
37
o1OpenAI · Closed
9.310%
38
Qwen3 235B 2507 (Reasoning)Alibaba · Open weight
8.481%
39
GPT-5 nanoOpenAI · Closed
8.276%
40
Qwen3.5 FlashAlibaba · Closed
6.207%
41
Claude Haiku 4.5Anthropic · Closed
5.903%
42
GPT-4.1OpenAI · Closed
5.517%
43
Gemini 2.5 FlashGoogle · Closed
4.844%
44
GPT-4.1 miniOpenAI · Closed
4.483%
45
GLM-4.6Z.AI · Open weight
3.819%
46
Grok 3 [Beta]xAI · Closed
3.793%
47
GLM-4.7Z.AI · Open weight
2.439%
48
Claude 3.5 SonnetAnthropic · Closed
2.069%
49
DeepSeek V3DeepSeek · Open weight
1.724%
50
GPT-4.1 nanoOpenAI · Closed
1.034%
51
Llama 4 MaverickMeta · Open weight
0.690%
52
GPT-4oOpenAI · Closed
0.345%
53
Llama 4 ScoutMeta · Open weight
0.000%

According to BenchLM.ai, GPT-5.6 Sol leads the FrontierMath v2 (Tiers 1-3) benchmark with a score of 89.000%, followed by GPT-5.6 Terra (84.900%) and GPT-5.6 Luna (78.600%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

53 models have been evaluated on FrontierMath v2 (Tiers 1-3). The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, FrontierMath v2 (Tiers 1-3) contributes 30% of the category score, so strong performance here directly affects a model's overall ranking.

About FrontierMath v2 (Tiers 1-3)

Year

2026

Tasks

295 private advanced mathematics problems

Format

Python-enabled iterative mathematical problem solving

Difficulty

From olympiad-plus to early research mathematics

After Epoch AI's June 2026 correction, the FrontierMath v2 private core contains 295 Tiers 1-3 problems. BenchLM selects the highest published reasoning-effort result for each model and keeps this core score distinct from Tier 4.

BenchLM freshness & provenance

Version

FrontierMath v2 (Tiers 1-3) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does FrontierMath v2 (Tiers 1-3) measure?

Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.

Which model scores highest on FrontierMath v2 (Tiers 1-3)?

GPT-5.6 Sol by OpenAI currently leads with a score of 89.000% on FrontierMath v2 (Tiers 1-3).

How many models are evaluated on FrontierMath v2 (Tiers 1-3)?

53 AI models have been evaluated on FrontierMath v2 (Tiers 1-3) on BenchLM.

Last updated: July 23, 2026 · BenchLM version FrontierMath v2 (Tiers 1-3) 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.