Skip to main content

Benchmark profile

Harvard-MIT Mathematics Tournament February 2026 (HMMT Feb 2026)

A February 2026 HMMT slice used in newer frontier-model math comparisons.

Data verified

Top models on HMMT Feb 2026 — July 23, 2026

As of July 23, 2026, Qwen3.7 Max leads the HMMT Feb 2026 leaderboard with 97.1% , followed by DeepSeek V4 Pro (Max) (95.2%) and DeepSeek V4 Flash (Max) (94.8%).

21 modelsMath25% of category scoreCurrentUpdated July 23, 2026

Leaderboard (21 models)

Score
1
Qwen3.7 MaxAlibaba · Closed
97.1%
2
DeepSeek V4 Pro (Max)DeepSeek · Open weight
95.2%
3
DeepSeek V4 Flash (Max)DeepSeek · Open weight
94.8%
4
DeepSeek V4 Pro (High)DeepSeek · Open weight
94.0%
5
Qwen3.7 PlusAlibaba · Closed
92.9%
6
Kimi K2.6Moonshot AI · Open weight
92.7%
7
GLM-5.2Z.AI · Open weight
92.5%
8
DeepSeek V4 Flash (High)DeepSeek · Open weight
91.9%
9
Qwen3.5 397BAlibaba · Open weight
87.9%
10
Qwen3.6 PlusAlibaba · Closed
87.8%
11
Kimi K2.5Moonshot AI · Open weight
87.1%
12
GLM-5Z.AI · Open weight
86.4%
13
Claude Opus 4.5Anthropic · Closed
85.3%
14
MAI-Thinking-1Microsoft · Closed
84.9%
15
Qwen3.6-27BAlibaba · Open weight
84.3%
16
Qwen3.6-35B-A3BAlibaba · Open weight
83.6%
17
GLM-5.1Z.AI · Open weight
82.6%
18
ZAYA1-8BZyphra · Open weight
71.6%
19
DeepSeek V4 FlashDeepSeek · Open weight
40.8%
20
DeepSeek V4 ProDeepSeek · Open weight
31.7%
21
MiniCPM5-1BOpenBMB · Open weight
25.8%

According to BenchLM.ai, Qwen3.7 Max leads the HMMT Feb 2026 benchmark with a score of 97.1%, followed by DeepSeek V4 Pro (Max) (95.2%) and DeepSeek V4 Flash (Max) (94.8%). The top models are clustered within 2.3 points, suggesting this benchmark is nearing saturation for frontier models.

21 models have been evaluated on HMMT Feb 2026. The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, HMMT Feb 2026 contributes 25% of the category score, so strong performance here directly affects a model's overall ranking.

About HMMT Feb 2026

Year

2026

Tasks

Competition math problems

Format

Contest mathematics

Difficulty

Olympiad-style mathematics

HMMT February 2026 matters because small score deltas at the frontier often depend on which contest set is used. BenchLM keeps this newer slice distinct from older HMMT summary rows.

BenchLM freshness & provenance

Version

HMMT Feb 2026 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does HMMT Feb 2026 measure?

A February 2026 HMMT slice used in newer frontier-model math comparisons.

Which model scores highest on HMMT Feb 2026?

Qwen3.7 Max by Alibaba currently leads with a score of 97.1% on HMMT Feb 2026.

How many models are evaluated on HMMT Feb 2026?

21 AI models have been evaluated on HMMT Feb 2026 on BenchLM.

Last updated: July 23, 2026 · BenchLM version HMMT Feb 2026 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.