Benchmark profile
LiveCodeBench v6
LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.
Data verifiedHow to read this leaderboard
Editorial review by Glevd · 2026-07-15
Compare v6 rows only when the date window, code-generation scenario, pass@k or average-at-k metric, sampling count, temperature, and execution policy match. A shared v6 label does not guarantee the rest of the setup is controlled.
Operator receipt: 10 sourced rows are currently displayable on this page; the leading published row is Sakana Fugu-Ultra at 93.2%.
Honest limit: Provider tables can use different v6 windows and generation settings. The route is a sourced release ledger, not a BenchLM rerun. LiveCodeBench still measures contest-style code generation rather than repository navigation, patch review, or regression safety.
Benchmark score on LiveCodeBench v6 — July 23, 2026
BenchLM mirrors the published score view for LiveCodeBench v6. Sakana Fugu-Ultra leads the public snapshot at 93.2% , followed by Sakana Fugu (92.9%) and Kimi K2.6 (89.6%). BenchLM does not use these results to rank models overall.
Sakana Fugu-Ultra
Sakana AI
sakana-fugu-ultra
Sakana Fugu
Sakana AI
sakana-fugu
Kimi K2.6
Moonshot AI
kimi-2-6
Benchmark score table (10 models)
ScoreThe published LiveCodeBench v6 snapshot places Sakana Fugu-Ultra first at 93.2%. The third row is 3.6 points behind. The broader top-10 range is 59.7 points, so the table still separates the published systems.
10 models have been evaluated on LiveCodeBench v6. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. LiveCodeBench v6 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About LiveCodeBench v6
Year
2026
Tasks
Fresh programming problems
Format
Provider-published v6 competitive programming results
Difficulty
Competitive programming level
Providers often publish a specific release or date window instead of the rolling aggregate. This route contains rows explicitly labeled v6 by their sources and excludes them from the weighted legacy lane.
BenchLM freshness & provenance
Version
LiveCodeBench v6 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does LiveCodeBench v6 measure?
LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.
Which model scores highest on LiveCodeBench v6?
Sakana Fugu-Ultra by Sakana AI currently leads with a score of 93.2% on LiveCodeBench v6.
How many models are evaluated on LiveCodeBench v6?
10 AI models have been evaluated on LiveCodeBench v6 on BenchLM.
Compare Top Models on LiveCodeBench v6
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.