Skip to main content

HLE (Humanity's Last Exam): The Hardest Benchmark

Humanity's Last Exam is crowdsourced from thousands of domain experts and designed to probe the absolute frontier of AI. Top models still top out under 65%. Here's why HLE matters.

Published
Last updated
Reading time
10 min
External sources
0
Tags: benchmarks, knowledge, hle, explainerData and scoring methodology
In this article5 sections

Looking for current scores? The live HLE leaderboard owns the ranking. This post explains the benchmark; the July ceiling analysis records how the 50% milestone fell.

HLE remains one of the hardest public AI benchmarks. The top displayable row is still below 65%, and the table spans frontier systems down to single-digit results. That spread is useful, but only when the protocol label travels with the score: a tool-assisted run and a closed-book run do not measure the same operating condition.

With MMLU and even GPQA approaching saturation, HLE remains the clearest measure of where frontier AI actually stands, and where it falls short.

What makes HLE different

HLE was crowdsourced from thousands of domain experts worldwide, organized by the Center for AI Safety and Scale AI. The questions are designed to test frontier-level knowledge (questions that even specialists find difficult), cover cutting-edge domains from advanced mathematics to theoretical physics and philosophy, resist memorization through novel expert-crafted questions not found in training data, and scale with AI progress so the benchmark stays challenging as models improve.

This isn't a test of whether a model can recall facts. It's a test of whether a model can reason at the level of the world's top researchers.

How questions are sourced

HLE's question creation process is unprecedented in scale. Over 3,000 domain experts from top universities and research institutions contributed questions. Each question goes through multiple validation rounds:

  1. Expert creates a question in their area of specialization, often at the frontier of their field
  2. Other experts verify the answer is correct and the question is appropriately difficult
  3. Difficulty calibration ensures questions require genuine expertise, not just encyclopedia knowledge
  4. Format standardization converts questions into consistent multiple-choice or short-answer formats

The result is a benchmark that probes knowledge most humans (even highly educated ones) simply don't have.

Current scores

HLE has substantial spread across 44 displayable sourced rows:

Table 1
Rank Model Score Evidence
1 Claude Mythos 5 64.5% Provider exact
2 Muse Spark 1.1 62.1% Provider exact
3 GPT-5.4 Pro 58.7% Provider exact
4 Claude Opus 4.8 57.9% Provider exact
5 Claude Sonnet 5 57.4% Provider exact

Full leaderboard: HLE scores

The table is a receipt, not a protocol-normalized race. Claude Mythos 5's attached source explicitly says “with tools”; other providers disclose different or less specific evaluation conditions. Use the gaps to shortlist models, then read the row provenance before treating a small difference as decisive.

Why the scores are so low

Even the best model scores below 65%, and most of the field is still under 50%. This tells us something important: current AI models have genuine limitations in deep expert reasoning. They're excellent at processing known information but still struggle with questions that require true expert-level insight.

Several factors contribute to the low scores:

Knowledge recency

Many HLE questions reference findings published after a model's training cutoff. A question about a 2025 theorem proof or a recent experimental result can't be answered from training data alone; it requires genuine reasoning about unfamiliar material.

Depth vs. breadth

Models trained on internet-scale data have extraordinary breadth. They know something about almost everything. But HLE tests depth: the kind of expertise that takes years of focused study in a narrow field. Current models are wide but not always deep enough.

Multi-step expert reasoning

The hardest HLE questions require chaining multiple pieces of specialist knowledge together. A physics question might require combining quantum field theory with statistical mechanics in a way that even PhD students find challenging. Models often get the individual pieces right but fail to connect them.

HLE vs other knowledge benchmarks

Table 2
Benchmark Top Score Score Spread Saturation Risk Best For
MMLU 93 70-93 High General knowledge baseline
MMLU-Pro 87 50-87 Medium Harder multiple choice
GPQA 97 80-97 High PhD-level science (3 domains)
SuperGPQA 95 55-95 Medium PhD-level (285 domains)
HLE 64.5 10-65 Very low Frontier reasoning

HLE is the only major knowledge benchmark still this far from its ceiling: the 50% mark was only crossed in 2026, and most models remain below it. This means it will remain useful for differentiating models for years.

When to use HLE for evaluation

HLE is most useful for comparing frontier models (it's the best benchmark for seeing real differences between GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro), for tracking AI progress over time since scores are far from saturation, and for assessing deep reasoning where PhD-level scientific knowledge is the requirement. Reasoning models tend to show larger HLE improvements than non-reasoning models, so it's also a good lens on whether chain-of-thought is earning its latency.

HLE is less useful for evaluating mid-tier models (most score in single digits) or for predicting performance on typical business tasks.

For calibration: human experts score about 74% on HLE questions in their own domain of expertise; outside their specialty, experts score much lower, sometimes below the models. Progress has been rapid: early 2025 frontier models scored around 10-15%, and by early 2026 scores have tripled to the mid-40s. The benchmark was designed to remain relevant as AI improves, and it's delivering on that promise.

All of which makes HLE the one score worth re-checking every quarter: the single most important benchmark for tracking AI progress at the frontier. While other knowledge tests have become checkboxes, this one still has an open ceiling; the 50% mark finally fell in 2026 with the Claude 5 family's tool-assisted runs, and the next milestones will show up here first.

See all models on the full leaderboard · Knowledge rankings


Reader questions

Frequently asked questions

01What is HLE (Humanity's Last Exam)?

HLE (Humanity's Last Exam) is an expert-authored AI benchmark with 3,000+ questions across advanced mathematics, science, humanities, and other specialist fields. It remains far from saturated: the top displayable score on BenchLM's July 14 snapshot is 64.5 from Claude Mythos 5 with tools.

02What score does GPT-5.4 get on HLE?

GPT-5.4 scores 52.1 on the currently attached HLE row. Claude Mythos 5 leads the July 14 displayable table at 64.5 with tools, followed by Muse Spark 1.1 at 62.1 and GPT-5.4 Pro at 58.7. Check row provenance before comparing protocols.

03Why do top AI models score so low on HLE?

HLE questions are at the frontier of human knowledge — many reference findings from after the model's training cutoff, require chaining multiple pieces of specialist knowledge, or test novel problem structures the model hasn't been trained to handle. Even human domain experts score only about 74% on questions in their own specialty. AI models have extraordinary breadth but not always the depth needed for these questions.

04How is HLE different from GPQA and MMLU?

MMLU covers 57 general subjects — frontier models score 97-99% (saturated). GPQA Diamond tests 3 science domains with PhD-level questions — top models score 95-97% (approaching saturation). HLE covers 100+ domains at the frontier of human knowledge — top models score 40-65% at best, with no sign of saturation. HLE provides the most differentiation between frontier models.

05Is HLE a good benchmark for comparing AI models?

HLE is the best public benchmark for comparing frontier models on deep reasoning. The 12-point gap between GPT-5.4 and Gemini 3.1 Pro would be invisible on MMLU-Pro (where they're within a point). However, HLE is less useful for mid-tier models, which mostly score in single digits with little variance between them.

Share or save

Share on XShare on LinkedIn

Keep reading

All research

New models drop every week. Join 2,000+ readers for one email a week on what moved, why, and what still needs proof.