Skip to main content

Benchmark profile

FrontierMath v2 Tier 4 (FrontierMath v2 (Tier 4))

Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.

Data verified

Top models on FrontierMath v2 (Tier 4) — July 23, 2026

As of July 23, 2026, GPT-5.6 Sol leads the FrontierMath v2 (Tier 4) leaderboard with 83.000% , followed by GPT-5.6 Terra (68.300%) and GPT-5.6 Luna (58.500%).

47 modelsMath10% of category scoreCurrentUpdated July 23, 2026

Leaderboard (47 models)

Score
1
GPT-5.6 SolOpenAI · Closed
83.000%
2
GPT-5.6 TerraOpenAI · Closed
68.300%
3
GPT-5.6 LunaOpenAI · Closed
58.500%
4
GPT-5.5 ProOpenAI · Closed
39.600%
5
GPT-5.4 ProOpenAI · Closed
37.500%
6
GPT-5.5OpenAI · Closed
35.400%
7
GPT-5.2 ProOpenAI · Closed
31.300%
8
Claude Opus 4.8Anthropic · Closed
31.250%
9
GPT-5.4OpenAI · Closed
27.100%
10
Claude Opus 4.7Anthropic · Closed
22.917%
11
Claude Opus 4.6Anthropic · Closed
22.900%
12
GPT-5.2OpenAI · Closed
18.800%
13
Gemini 3 ProGoogle · Closed
18.750%
14
Gemini 3.1 ProGoogle · Closed
16.700%
15
Muse SparkMeta · Closed
14.600%
16
Gemini 3.5 FlashGoogle · Closed
14.583%
17
Kimi K2.6Moonshot AI · Open weight
14.580%
18
GLM-5.1Z.AI · Open weight
12.500%
19
GPT-5.1OpenAI · Closed
12.500%
20
Qwen3.6 PlusAlibaba · Closed
8.333%
21
Claude Sonnet 4.6Anthropic · Closed
8.300%
22
GPT-5.4 nanoOpenAI · Closed
6.250%
23
o4-mini (high)OpenAI · Closed
6.250%
24
GPT-5 miniOpenAI · Closed
6.250%
25
Kimi K2.5Moonshot AI · Open weight
4.200%
26
Claude Opus 4.5Anthropic · Closed
4.167%
27
Qwen 3.6 Max (preview)Alibaba · Closed
4.167%
28
Gemini 3 FlashGoogle · Closed
4.167%
29
Claude Sonnet 4.5Anthropic · Closed
4.167%
30
Gemini 2.5 FlashGoogle · Closed
4.167%
31
Gemini 2.5 ProGoogle · Closed
4.167%
32
GLM-4.6Z.AI · Open weight
2.128%
33
GLM-5Z.AI · Open weight
2.100%
34
DeepSeek V3.2DeepSeek · Open weight
2.100%
35
Claude Haiku 4.5Anthropic · Closed
2.083%
36
Qwen3.5 PlusAlibaba · Closed
2.083%
37
Grok 4xAI · Closed
2.083%
38
o3OpenAI · Closed
2.083%
39
GPT-5.4 miniOpenAI · Closed
2.080%
40
GLM-4.7Z.AI · Open weight
0.000%
41
Kimi K2Moonshot AI · Closed
0.000%
42
GPT-4.1OpenAI · Closed
0.000%
43
Qwen3 235B 2507 (Reasoning)Alibaba · Open weight
0.000%
44
Qwen3.5 FlashAlibaba · Closed
0.000%
45
Grok 3 [Beta]xAI · Closed
0.000%
46
Claude 3.5 SonnetAnthropic · Closed
0.000%
47
GPT-5 nanoOpenAI · Closed
0.000%

According to BenchLM.ai, GPT-5.6 Sol leads the FrontierMath v2 (Tier 4) benchmark with a score of 83.000%, followed by GPT-5.6 Terra (68.300%) and GPT-5.6 Luna (58.500%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

47 models have been evaluated on FrontierMath v2 (Tier 4). The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, FrontierMath v2 (Tier 4) contributes 10% of the category score, so strong performance here directly affects a model's overall ranking.

About FrontierMath v2 (Tier 4)

Year

2026

Tasks

43 private extreme-difficulty mathematics problems

Format

Python-enabled iterative mathematical problem solving

Difficulty

Research-level mathematics requiring hours or days of expert work

The v2 private Tier 4 set contains 43 problems. Its smaller sample makes individual scores noisier than Tiers 1-3, so it is a separate ranking factor with lower weight rather than being averaged into the core score.

BenchLM freshness & provenance

Version

FrontierMath v2 (Tier 4) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does FrontierMath v2 (Tier 4) measure?

Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.

Which model scores highest on FrontierMath v2 (Tier 4)?

GPT-5.6 Sol by OpenAI currently leads with a score of 83.000% on FrontierMath v2 (Tier 4).

How many models are evaluated on FrontierMath v2 (Tier 4)?

47 AI models have been evaluated on FrontierMath v2 (Tier 4) on BenchLM.

Last updated: July 23, 2026 · BenchLM version FrontierMath v2 (Tier 4) 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.