Frontier AI models now score 95-99% on AIME and HMMT; competition math is effectively solved. The top 5 models are within 2 points of each other on both benchmarks. For comparing frontier models on math in 2026, BRUMO and MATH-500 provide more signal. AIME and HMMT remain useful as display benchmarks and floor checks for mid-tier models, but we no longer weight them into the math score.
The American Invitational Mathematics Examination (AIME) and Harvard-MIT Mathematics Tournament (HMMT) are prestigious math competitions designed for the most talented high school students. They've become standard AI benchmarks, and the results are striking.
Frontier models now score 95-99% on these competitions. Competition-level math is, for practical purposes, solved by AI.
What AIME and HMMT test
AIME is a 15-question, 3-hour examination. Each answer is an integer from 000 to 999. The problems require creative mathematical insight across algebra, geometry, number theory, and combinatorics.
In human competition, qualifying for AIME puts a student in the top ~5% nationally. A perfect score is exceptionally rare; in most years, fewer than a handful of students achieve it.
What makes AIME challenging is that problems rarely require advanced mathematical knowledge. Instead, they demand creative problem-solving: seeing non-obvious connections, applying techniques in novel ways, and constructing multi-step proofs. This is precisely why AIME became popular as an AI benchmark: it tests genuine mathematical reasoning.
We track three years: AIME 2023, AIME 2024, and AIME 2025. Tracking multiple years helps detect whether models memorized specific problem sets or have generalizable math ability.
HMMT is the harder sibling. It is hosted jointly by Harvard and MIT and is one of the most competitive high school math tournaments in the US. Problems span algebra, geometry, combinatorics, and number theory at a difficulty comparable to or exceeding AIME, with more emphasis on proof-like reasoning and multi-step deductions.
We track the same three years here: HMMT 2023, HMMT 2024, and HMMT 2025.
Current scores
| Model | AIME 2025 | HMMT 2025 |
|---|---|---|
| GPT-5.4 | 99 | 97 |
| GPT-5.3 Codex | 98 | 96 |
| Claude Opus 4.6 | 98 | 96 |
| Grok 4.1 | 98 | 96 |
The top models are all above 95 on AIME 2025 and above 90 on HMMT 2025. The gaps between models are just 1-2 points, within noise range.
It did not start this way. Competition math benchmarks followed a predictable arc. In 2023, the best models scored around 50-60% on AIME. By mid-2024, reasoning-enhanced models pushed scores into the 80s. By early 2025, scores crossed 90%, and by 2026, the benchmark is effectively saturated.
This rapid progression illustrates a broader pattern: once models develop the right reasoning capabilities, scores compress quickly. The 50-to-95 jump happened in roughly 18 months.
Year-over-year contamination risk
A model that scores 98 on AIME 2023 but only 85 on AIME 2025 might have memorized older problems. We track three consecutive years so you can spot exactly this pattern. In practice, frontier models score consistently across all three years, suggesting genuine mathematical ability.
What "solved" means
When we say competition math is "solved," we mean AI models can reliably answer these problems at or above the level of the best human competitors. The 1-2 point differences between frontier models aren't meaningful.
However, some caveats apply:
- Chain-of-thought helps enormously: Without it, even frontier models score 10-20 points lower on AIME. The reasoning process itself is critical.
- Formatting matters: Models sometimes produce correct reasoning but format the final integer answer incorrectly.
- Verification is different from generation: Models scoring 98 can solve problems but can't always explain why their approach works at the level a human mathematician would.
Benchmarks that still differentiate
For meaningful separation between frontier models on math:
- BRUMO 2025: Bulgarian Mathematical Olympiad, slightly more separation
- MATH-500: broader difficulty range, more variance in mid-tier models
- HLE: includes advanced math at frontier difficulty (top models score 10-46%)
→ See all math models ranked · Full leaderboard
Competition math benchmarks served their purpose during 2023-2024 when they could still separate models. In 2026, AIME and HMMT are floor checks; that is why we no longer weight them into the math score. When you compare frontier models on math, start from BRUMO 2025 and MATH-500 instead.
Reader questions
Frequently asked questions
01What is AIME and how is it used to benchmark AI?
AIME (American Invitational Mathematics Examination) is a prestigious US high school math competition with 15 integer-answer problems. It requires creative mathematical insight across algebra, geometry, number theory, and combinatorics. AI models are tested on AIME because it demands genuine mathematical reasoning rather than rote knowledge — the same property that made it hard for humans makes it a good AI benchmark.
02What do AI models score on AIME 2025?
As of May 2026, GPT-5.4 scores 99 on AIME 2025, while GPT-5.3 Codex, Claude Opus 4.6, and Grok 4.1 score 98. The top frontier models are all above 95 — competition math is effectively solved at the frontier.
03Is AIME still a useful AI benchmark in 2026?
AIME is largely saturated for frontier model comparison — the top 5 models are all above 95, with only 1-2 point differences that fall within noise range. BRUMO 2025 and MATH-500 still show more separation between frontier models. AIME 2025 and HMMT 2025 remain useful as display benchmarks and floor checks for mid-tier models, but they no longer factor into our weighted math score.
04What is HMMT and how does it compare to AIME?
HMMT (Harvard-MIT Mathematics Tournament) is a high school math competition jointly hosted by Harvard and MIT. Problems are comparable to or harder than AIME, with emphasis on proof-like reasoning and multi-step deductions. Frontier AI models score 96-97 on HMMT 2025, slightly lower than AIME 2025, suggesting HMMT still provides marginally more discrimination.
05What math benchmarks still differentiate frontier AI models in 2026?
BRUMO 2025 (Bulgarian Mathematical Olympiad) and MATH-500 still show meaningful separation between frontier models. All AIME and HMMT variants are now display-only in our rankings because they are too saturated at the frontier. For the hardest math reasoning, HLE (Humanity's Last Exam) includes advanced mathematical content where top models score only 10-46%.
Continue with live BenchLM data
Share or save