Skip to main content

Benchmark profile

LiveCodeBench v6

LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.

Data verified

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Compare v6 rows only when the date window, code-generation scenario, pass@k or average-at-k metric, sampling count, temperature, and execution policy match. A shared v6 label does not guarantee the rest of the setup is controlled.

Operator receipt: 10 sourced rows are currently displayable on this page; the leading published row is Sakana Fugu-Ultra at 93.2%.

Honest limit: Provider tables can use different v6 windows and generation settings. The route is a sourced release ledger, not a BenchLM rerun. LiveCodeBench still measures contest-style code generation rather than repository navigation, patch review, or regression safety.

Benchmark score on LiveCodeBench v6 — July 23, 2026

BenchLM mirrors the published score view for LiveCodeBench v6. Sakana Fugu-Ultra leads the public snapshot at 93.2% , followed by Sakana Fugu (92.9%) and Kimi K2.6 (89.6%). BenchLM does not use these results to rank models overall.

10 modelsCodingCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (10 models)

Score
1
Sakana Fugu-UltraSakana AI · Closed
93.2%
2
Sakana FuguSakana AI · Closed
92.9%
3
Kimi K2.6Moonshot AI · Open weight
89.6%
4
Qwen3.6 PlusAlibaba · Closed
87.1%
5
Kimi K2.5Moonshot AI · Open weight
85.0%
6
Claude Opus 4.5Anthropic · Closed
84.8%
7
Qwen3.5 397BAlibaba · Open weight
83.6%
8
ZAYA1-8BZyphra · Open weight
65.8%
9
ZAYA1-74B-PreviewZyphra · Open weight
65.7%
10
MiniCPM5-1BOpenBMB · Open weight
33.5%

The published LiveCodeBench v6 snapshot places Sakana Fugu-Ultra first at 93.2%. The third row is 3.6 points behind. The broader top-10 range is 59.7 points, so the table still separates the published systems.

10 models have been evaluated on LiveCodeBench v6. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. LiveCodeBench v6 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About LiveCodeBench v6

Year

2026

Tasks

Fresh programming problems

Format

Provider-published v6 competitive programming results

Difficulty

Competitive programming level

Providers often publish a specific release or date window instead of the rolling aggregate. This route contains rows explicitly labeled v6 by their sources and excludes them from the weighted legacy lane.

BenchLM freshness & provenance

Version

LiveCodeBench v6 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does LiveCodeBench v6 measure?

LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.

Which model scores highest on LiveCodeBench v6?

Sakana Fugu-Ultra by Sakana AI currently leads with a score of 93.2% on LiveCodeBench v6.

How many models are evaluated on LiveCodeBench v6?

10 AI models have been evaluated on LiveCodeBench v6 on BenchLM.

Last updated: July 23, 2026 · BenchLM version LiveCodeBench v6 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.