The July 14 registry contains 296 benchmark definitions across 10 categories. That count is the operator receipt behind this review; it replaces the 290 in the previous title rather than merely advancing the date.
Claude Mythos 5 leads overall, coding, and agentic work. Claude Fable 5 is second on all three. GPT-5.6 Sol is third on all three, though its overall and coding rows carry Estimated evidence. The repetition is the important finding in July: three independent ranking views now tell the same broad story.
The method changed too. BenchLM no longer treats an unreported benchmark as a zero or forces every model through an identical publication checklist. BenchAlign v5 estimates comparable capability, preserves uncertainty, and tells readers whether the result is Supported or Estimated.
Top models overall
| Rank | Model | Type | License | Score | Evidence |
|---|---|---|---|---|---|
| 1 | Claude Mythos 5 | Reasoning | Proprietary | 83.9 | Supported |
| 2 | Claude Fable 5 | Reasoning | Proprietary | 83.7 | Supported |
| 3 | GPT-5.6 Sol | Reasoning | Proprietary | 82 | Supported |
| 4 | Kimi K3 | Reasoning | Pending | 81 | Supported |
| 5 | Claude Opus 4.8 | Reasoning | Proprietary | 78.3 | Supported |
| 6 | Muse Spark 1.1 | Reasoning | Proprietary | 77.4 | Supported |
| 7 | Grok 4.5 | Reasoning | Proprietary | 76.7 | Supported |
| 8 | GPT-5.4 | Reasoning | Proprietary | 74.2 | Supported |
| 9 | GPT-5.5 | Reasoning | Proprietary | 73.5 | Estimated |
| 10 | Qwen3.7 Max | Reasoning | Proprietary | 72.8 | Supported |
Mythos is the capability leader, but it is restricted. Fable is the leading generally available row. GPT-5.6 Sol is the closest non-Anthropic challenger, and its Estimated overall label matters: the position is useful, but less settled than a Supported row with comparable evidence breadth.
Claude Opus 4.8 remains fourth overall at 77.8. That corrects another misleading old story. Opus 4.6 is not Anthropic's frontier representative, and a comparison that leaves out Opus 4.8, Fable, or Mythos is historical by definition.
Coding leaders
| Rank | Model | Type | License | Score | Evidence |
|---|---|---|---|---|---|
| 1 | Claude Mythos 5 | Reasoning | Proprietary | 80.8 | Supported |
| 2 | Claude Fable 5 | Reasoning | Proprietary | 80.6 | Supported |
| 3 | GPT-5.6 Sol | Reasoning | Proprietary | 78.9 | Supported |
| 4 | Kimi K3 | Reasoning | Pending | 78 | Supported |
| 5 | GPT-5.6 Luna | Reasoning | Proprietary | 72.6 | Supported |
| 6 | GPT-5.5 | Reasoning | Proprietary | 71.4 | Supported |
| 7 | Claude Opus 4.8 | Reasoning | Proprietary | 70.6 | Supported |
| 8 | Claude Opus 4.7 | Non-Reasoning | Proprietary | 68.4 | Supported |
| 9 | Claude Sonnet 5 | Reasoning | Proprietary | 68.4 | Supported |
| 10 | GPT-5.6 Terra | Reasoning | Proprietary | 66 | Supported |
The coding order is not derived from one SWE-bench result. BenchAlign combines the coding evidence, normalizes each benchmark against its competitive field, and limits the influence of overlapping or repeated sources. That avoids equating 90 on an easier benchmark with 90 on a harder one.
The first four rows deserve different readings. Mythos and Fable have Supported evidence and lead. GPT-5.6 Sol is third with Estimated evidence. Opus 4.8 is fourth with Supported evidence. A team choosing between Sol and Opus should not read a 0.59-point score gap as certainty; it should run both on its own repositories.
Agentic leaders
| Rank | Model | Type | License | Score | Evidence |
|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | Reasoning | Proprietary | 75.2 | Supported |
| 2 | GPT-5.6 Terra | Reasoning | Proprietary | 73.1 | Supported |
| 3 | Muse Spark 1.1 | Reasoning | Proprietary | 68.6 | Supported |
| 4 | Kimi K3 | Reasoning | Pending | 66.6 | Supported |
| 5 | Claude Opus 4.7 (Adaptive) | Reasoning | Proprietary | 65.9 | Supported |
| 6 | GPT-5.5 | Reasoning | Proprietary | 65.8 | Supported |
| 7 | Claude Mythos 5 | Reasoning | Proprietary | 65 | Supported |
| 8 | Claude Opus 4.8 | Reasoning | Proprietary | 64 | Supported |
| 9 | Claude Fable 5 | Reasoning | Proprietary | 63.9 | Supported |
| 10 | GPT-5.3 Codex | Reasoning | Proprietary | 61.2 | Estimated |
Agentic work is where the top three are best supported. Mythos, Fable, and Sol all carry Supported evidence. Their spread is only 1.6 points from first to third. The ranking is informative, but an operating decision should also measure intervention rate, tool-call recovery, latency, and the cost of a failed run.
Older rows such as o1 and o3 can still exist in the catalog without being sensible current recommendations. A benchmark database is also a historical record. The current rank should favor the current model generation when the evidence does.
Open-weight leaders
| Rank | Model | Type | License | Score | Evidence |
|---|---|---|---|---|---|
| 1 | MiniMax M3 | Non-Reasoning | Open Weight | 69.8 | Supported |
| 2 | GLM-5.1 | Reasoning | Open Weight | 67.7 | Supported |
| 3 | Inkling | Hybrid | Open Weight | 67.5 | Supported |
| 4 | GLM-5 | Non-Reasoning | Open Weight | 66.1 | Supported |
| 5 | MiniMax M2.7 | Non-Reasoning | Open Weight | 64.1 | Supported |
| 6 | GLM-5.2 | Reasoning | Open Weight | 64 | Estimated |
| 7 | GLM-4.7 | Reasoning | Open Weight | 61.2 | Supported |
| 8 | Gemma 4 31B | Reasoning | Open Weight | 61.1 | Supported |
| 9 | Qwen3.5-27B | Reasoning | Open Weight | 60.7 | Supported |
| 10 | DeepSeek V4 Pro | Non-Reasoning | Open Weight | 60.7 | Supported |
MiniMax M3 leads the open-weight overall table at 69.8. GLM-5.1 and GLM-5 follow. The gap from Mythos is 14.05 points, so broad open-weight parity has not arrived. The operational comparison is narrower than that number suggests when privacy, fine-tuning, data residency, or high-volume serving makes an API model unsuitable.
The open-weight coding leader is different from the broad leader. GLM-5.2 and Kimi K2.7 Code sit at the top of the current open coding view. That is why the site keeps overall, coding, and agentic surfaces separate: the best general model is not automatically the best specialist model.
What BenchAlign v5 changes
The old ranking had two bad options for sparse models. Ignore missing benchmarks and a model could rise on a small set of favorable disclosures. Convert missing benchmarks to zeros and a model could be punished for what its lab chose not to publish. Requiring a fixed coverage threshold avoided some false precision, but it also left strong new models unranked.
BenchAlign v5 uses six controls:
- Within-benchmark normalization. A score is judged against that benchmark's field, not against the raw percentage scale of another test.
- Source and benchmark de-duplication. Closely related evidence cannot vote repeatedly as though it were independent.
- Reliability weighting. Independent, reproducible, difficult-to-game sources carry more weight than weak or self-reported evidence.
- Uncertainty-aware estimation. Missing observations widen uncertainty instead of turning into automatic failure.
- Family consistency checks. Model-family information can stabilize sparse rows without forcing a newer or “Pro” label to win against contrary evidence.
- Supported and Estimated status. Readers can distinguish a well-supported rank from a useful but less certain estimate.
No method makes the publication incentives disappear. It makes them visible and reduces the reward for selective disclosure.
Which benchmarks still matter?
Benchmarks matter when they separate current systems, resemble a consequential task, resist contamination, and publish enough detail to audit. Hard coding and agentic evaluations remain useful because model failures are observable and the field still has spread. HLE, GPQA, and MMLU-Pro remain more informative than saturated legacy knowledge tests. Grounded document and multimodal evaluations matter for users whose work is not plain text.
That does not mean every useful benchmark belongs in one universal score. Math and multilingual pages can remain lenses for readers and search engines while the core rank focuses on the six capability pillars. Multimodal evidence should inform versatility and relevant use cases without punishing a text-only configuration for not accepting images.
The July verdict
The current broad order is Mythos, Fable, GPT-5.6 Sol, and Opus 4.8. MiniMax M3 leads open weight. Mythos also leads coding and agentic work, with Fable second and Sol third.
The more important change is that a model can now rank without publishing the exact same benchmark set as every competitor. Its uncertainty remains visible. That is a better answer than either excluding it or pretending its missing rows are failures.
→ Live leaderboard · Coding ranking · Agentic ranking · Open-weight guide
Reader questions
Frequently asked questions
01What is the best AI model in 2026?
Claude Mythos 5 leads BenchLM's July 2026 overall ranking at 83.85 with Supported evidence. Claude Fable 5 is second at 83.6, GPT-5.6 Sol is third at 79.3 with Estimated evidence, and Claude Opus 4.8 is fourth at 77.8. Mythos access is restricted, so Fable is the leading generally available row.
02Which LLM is best for coding in 2026?
Claude Mythos 5 leads the current coding ranking at 81.95, followed by Claude Fable 5 at 81.7. GPT-5.6 Sol is third at 74.08 with Estimated evidence, and Claude Opus 4.8 is fourth at 73.49 with Supported evidence.
03Which model is best for AI agents?
Claude Mythos 5 leads the agentic ranking at 77.09, Claude Fable 5 follows at 76.84, and GPT-5.6 Sol is third at 75.49. All three agentic rows have Supported evidence.
04What is the best open-weight model?
MiniMax M3 leads BenchLM's open-weight overall ranking at 69.8 with Supported evidence. GLM-5.1 follows at 67.76, then GLM-5 at 65.98.
05How does BenchLM handle missing benchmarks?
BenchAlign v5 estimates comparable capability from the evidence a model has, carries uncertainty forward, and labels each row Supported or Estimated. Missing results no longer become zeros, and models no longer disappear merely because a lab did not publish the same benchmark set as its competitors.
Share or save