Benchmark profile
Agents Last Exam (ALE-Bench)
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
How BenchLM shows ALE-Bench
BenchLM mirrors the Agents Last Exam full split from the public leaderboard API. The snapshot reports pass rate, partial average score, cost, token, and duration metadata across 152 professional workflow tasks for model plus agent-harness rows.
ALE-Bench is display only on BenchLM. Its rows combine a base model with an agent harness such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, or Gemini CLI, so BenchLM keeps the table separate from model-only rankings.
The Agent Showdown analysis adds domain and failure-mode context across 13 top-level domains. It also notes that Claude Code plus Fable 5 may include fallback to Opus 4.8 on refused tasks, so BenchLM preserves the official mixed-system row label instead of treating it as a pure base-model score.
Pass rate on ALE-Bench — June 2026 API snapshot
BenchLM mirrors the published pass rate view for ALE-Bench. codex (reasoning-xhigh) / GPT-5.6-Sol leads the public snapshot at 30.6% , followed by codex (reasoning-high) / GPT-5.6-Sol (30.6%) and codex (reasoning-max) / GPT-5.6-Sol (29.6%). BenchLM does not use these results to rank models overall.
codex (reasoning-xhigh) / GPT-5.6-Sol
codex
codex:reasoning-xhigh/GPT-5.6-Sol
codex (reasoning-high) / GPT-5.6-Sol
codex
codex:reasoning-high/GPT-5.6-Sol
codex (reasoning-max) / GPT-5.6-Sol
codex
codex:reasoning-max/GPT-5.6-Sol
Pass rate table (53 models)
ScoreThe published ALE-Bench snapshot places codex (reasoning-xhigh) / GPT-5.6-Sol first at 30.6%. The third row is 1.0 points behind. The broader top-10 range is 4.0 points, so many of the published results sit in a relatively narrow band.
53 models have been evaluated on ALE-Bench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. ALE-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About ALE-Bench
Year
2026
Tasks
152 ALE-V1 professional workflow tasks across 13 top-level domains
Format
Pass rate, partial-credit score, cost, token, and duration metadata
Difficulty
Real-world agentic workflows
BenchLM mirrors the public Agents Last Exam full leaderboard API as ALE-Bench and links the June 2026 Agent Showdown analysis for domain, cost, speed, and failure-mode context. Rows combine base models with agent harnesses such as Codex, OpenClaw, Claude Code, Droid, Cursor CLI, and Gemini CLI, so the table remains display-only. The source notes that Claude Code plus Fable 5 may include upstream fallback to Opus 4.8 on refused tasks.
BenchLM freshness & provenance
Version
ALE-Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does ALE-Bench measure?
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
Which model leads the published ALE-Bench snapshot?
codex (reasoning-xhigh) / GPT-5.6-Sol currently leads the published ALE-Bench snapshot with 30.6% pass rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on ALE-Bench?
53 AI models are included in BenchLM's mirrored ALE-Bench snapshot, based on the public leaderboard captured on June 2026 API snapshot.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.