Best LLM for Coding (July 2026): SWE-bench & LiveCodeBench Ranked
Claude Mythos 5 leads coding on BenchLM's July 2026 rankings with a score of 80.8, ahead of Claude Fable 5 (80.6) and GPT-5.6 Sol (78.9). Supported and Estimated labels show the evidence maturity behind each position.
Data refreshed:
Programming and software development
Decision lens: the score determines position; Supported and Estimated labels describe the evidence behind that position without removing sparsely reported models.
- Data refreshed
- July 23, 2026
- Ranked
- 122 of 290 models
- Supported / Estimated
- 44 / 78
- Weighted evidence
- 2 of 8 benchmarks
8 tracked benchmarks
HumanEval, SWE-bench Verified, LiveCodeBench, LiveCodeBench Pro, FLTEval, SWE-bench Pro, SWE-Rebench, SWE Multilingual, CursorBench, Multi-SWE Bench, VIBE-Pro, NL2Repo, Vibe Code Bench, React Native Evals, SWE-bench Verified*, Spider 2.0-Lite
Evidence set: HumanEval, SWE-bench Verified, LiveCodeBench, LiveCodeBench Pro, FLTEval, SWE-bench Pro, SWE-Rebench, SWE Multilingual, CursorBench, Multi-SWE Bench, VIBE-Pro, NL2Repo, Vibe Code Bench, React Native Evals, SWE-bench Verified*, Spider 2.0-Lite
Best Coding picks
BenchLM summaries for coding plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
SWE-bench Pro & LiveCodeBench Leaderboard
Primary score: BenchAlign coding score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Supported positions have diverse direct evidence. Estimated positions remain ranked but carry wider uncertainty.
Filters
1 | 80.8% | 80.8 | — | 95.5% | — | — | — | 80.3% | — | — | — | — | — | — | — | — | — | — |
2 | 80.5% | 80.55 | — | 95% | — | — | — | 80% | — | — | 70.5% | — | — | — | — | — | — | — |
3 | 78.9% | 78.89 | — | — | — | — | — | 64.6% | — | — | 67.2% | — | — | — | — | — | — | — |
4 | 78.0% | 78.02 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
5 | 72.6% | 72.56 | — | — | — | — | — | 62.7% | — | — | 61.1% | — | — | — | — | — | — | — |
6 | 71.4% | 71.41 | — | — | — | — | — | 58.6% | — | — | 58.4% | — | — | — | 69.85% | 84.7% | — | — |
7 | 70.6% | 70.61 | — | 88.6% | — | — | — | 69.2% | — | 84.4% | 62.3% | — | — | — | — | — | — | — |
8 | 68.4% | 68.39 | — | — | — | — | — | — | — | — | — | — | — | — | 71.00% | 82.8% | — | — |
9 | 68.4% | 68.38 | — | 85.2% | — | — | — | 63.2% | — | 78.3% | 61.5% | — | — | — | — | — | — | — |
10 | 66.0% | 65.99 | — | — | — | — | — | 63.4% | — | — | 64.9% | — | — | — | — | — | — | — |
11 | 65.1% | 65.12 | — | — | — | — | — | 61.5% | — | — | — | — | — | — | — | — | — | — |
| 65.1% | 65.09 | — | — | — | — | — | 62.1% | — | — | 55.0% | — | — | 48.9% | — | — | — | — | |
13 | 64.2% | 64.2 | — | 77.4% | — | 80.0% | — | 52.4% | — | — | — | — | — | — | 19.67% | — | — | — |
14 | 64.2% | 64.2 | — | 85% | — | — | — | 56.8% | 58.2% | — | — | — | — | — | 61.77% | — | — | — |
15 | 62.8% | 62.83 | — | 87.6% | — | — | — | 64.3% | — | — | — | — | — | — | — | — | — | — |
16 | 62.8% | 62.78 | — | — | — | — | — | — | — | — | — | — | — | — | 14.30% | — | — | — |
17 | 62.6% | 62.58 | — | 78% | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
18 | 62.5% | 62.49 | — | 80.8% | — | 70.7% | — | 53.4% | 65.3% | — | — | — | — | — | 57.57% | 84.1% | 75.6% | — |
19 | 62.0% | 61.98 | — | — | — | — | — | 55.1% | — | — | 48.8% | — | — | — | 48.68% | — | — | — |
20 | 61.9% | 61.93 | — | 80.9% | — | — | — | 57.1% | — | 77.5% | — | — | — | 43.2% | — | — | — | — |
21 | 60.7% | 60.72 | — | — | — | — | — | — | — | — | — | — | — | — | 53.50% | — | — | — |
22 | 59.6% | 59.61 | — | — | — | 87.5% | — | 57.7% | — | — | — | — | — | — | 67.42% | 85.3% | — | — |
23 | 59.6% | 59.56 | — | 79.6% | — | — | — | — | 60.7% | — | — | — | — | — | 51.48% | 80.6% | — | — |
24 | 59.4% | 59.38 | — | 77.8% | — | — | — | 55.1% | 62.8% | 73.3% | — | — | — | — | — | 74.8% | 72.8% | — |
25 | 59.1% | 59.1 | — | — | — | — | — | 64.7% | — | 78% | 66.7% | — | — | — | — | — | — | — |
Top AI Models for Coding — July 2026
As of July 2026, Claude Mythos 5 leads the BenchAlign coding leaderboard with a score of 80.8, followed by Claude Fable 5 (80.5) and GPT-5.6 Sol (78.9). BenchLM is currently showing 44 Supported and 78 Estimated models in this category.
Ranks #1 on the current coding board with a Supported evidence label.
Ranks #2 on the current coding board with a Supported evidence label.
Ranks #3 on the current coding board with a Supported evidence label.
What changed
Claude Mythos 5 ranks #1 at 80.8 with a Supported evidence label.
Claude Fable 5 ranks #2 at 80.5 with a Supported evidence label.
GPT-5.6 Sol ranks #3 at 78.9 with a Supported evidence label.
Top models by benchmark
Real-world GitHub issues from popular Python repos, human-verified subset(16% of category score)
Score in Context
What these scores mean
BenchAlign places direct benchmarks and independent external signals on a common calibrated scale. The score is relative to the current evidence universe; it is not a raw percentage from any single test.
Known limitations
Estimated rows have less diverse direct evidence and wider uncertainty. They remain ranked so a newly released model is not treated as weak merely because fewer benchmark publishers have evaluated it.
How we weight
This lens combines category-relevant external evidence with admitted benchmark protocols. Evidence sources are calibrated for difficulty before aggregation, and no generated benchmark row contributes to the score.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| SWE-bench Pro | 50% | Weighted | Frontier real-world SE tasks |
| LiveCodeBench | 50% | Weighted | Contamination-free competitive programming |
| SWE-Rebench | — | Display only | Fresh rolling-window GitHub issues |
| ProgramBench | — | Display only | Cleanroom full-program reconstruction |
| SWE-bench Verified | — | Display only | Historical baseline, superseded by Pro |
| FLTEval | — | Display only | Lean 4 proof engineering, sparse coverage |
| React Native Evals | — | Display only | Framework-specific mobile app engineering |
| HumanEval | — | Display only | Saturated by frontier models |
About Coding Benchmarks
Python programming problems with test cases
Common questions
Which LLM is best for coding?
The model in the #1 row of the live leaderboard above is BenchLM's current best LLM for coding. Rankings are recomputed on every data refresh from a weighted blend of SWE-bench Pro (real GitHub issues) and LiveCodeBench (contamination-resistant competitive programming), so the answer box at the top of this page always names the current leader and its score rather than a snapshot that can go stale.
What is the best LLM for coding right now?
Right now the top three coding models are shown in the answer box and Top ranked panel on this page, updated with each leaderboard refresh. The current leaders separate themselves on SWE-bench Pro, the hardest widely run software engineering benchmark, where a few points of difference typically decide whether a model can resolve a multi-file GitHub issue end to end. Check the live table for today's exact ordering and scores.
What is the best free LLM for coding?
Free usually means one of two things: a free chat tier for a proprietary model, or open weights you can download and run yourself. Most frontier coding models offer rate-limited free tiers in their chat apps, while open-weight models cost nothing to self-host beyond compute. For the strongest no-cost option, start with the highest-ranked open-weight model on this leaderboard, then compare it in our best open-source LLM ranking.
What is the best open source LLM for coding?
The best open-source coding model is the highest-ranked row marked Open Weight on the leaderboard above. Open-weight models now sit within a few points of the proprietary frontier on SWE-bench Pro and LiveCodeBench, and can be self-hosted, fine-tuned, and run without per-token API pricing. Our best open-source LLM page ranks them across every category and is the fastest way to find the current open-weight leader.
How do you benchmark an LLM's coding ability?
By executing the model's code, not by grading it subjectively. SWE-bench Pro hands models real GitHub issues and counts a task solved only when the generated patch passes the repository's own test suite. LiveCodeBench scores freshly published competitive-programming problems to rule out training-data contamination. BenchLM weights those two benchmarks equally for the coding score, and tracks display-only signals such as DeepSWE, which measures long-horizon software engineering with agent harnesses.
Coding leaderboard updates
Know which model codes best — before your team picks the wrong one.
One email each week. Unsubscribe anytime.
Related
Best LLMs Overall
Top models ranked across all benchmark categories.
Best Open-Weight Models
Top open-source models for code generation and debugging.
Agentic Benchmarks
How models perform on autonomous coding agent tasks.
AI Cost Calculator
Compare pricing across models for coding workloads.