Skip to main content

Model comparison

Sakana Fugu vs Sakana Fugu-Ultra

Data verified

Head-to-head evidence from 11 shared benchmark results across 5 categories. Overall scores shown here use the public BenchAlign v5 ranking lane.

Sibling matchup inside the Sakana Fugu family.

Sakana AI
N/A
No comparison
N/A
0 category wins4 category wins

Evidence parity. Sakana Fugu and Sakana Fugu-Ultra share 11 comparable benchmark results. 5 of 8 categories are comparable. 0 results are unique to Sakana Fugu; 0 to Sakana Fugu-Ultra.

Updated July 23, 2026
Shared results
11
Sakana Fugu only
0
Sakana Fugu-Ultra only
0
Comparable categories
5 / 8

Sakana Fugu makes more sense if you want this variant’s specific profile, while Sakana Fugu-Ultra is the cleaner fit if reasoning is the priority.

Confidence note. This is a partial-evidence comparison with 11 shared benchmark results across 5 evidence categories; 5 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.

Why this result

Sakana Fugu and Sakana Fugu-Ultra sit in the same Sakana Fugu family. This page is less about two unrelated model lineages and more about how the siblings trade off on benchmark shape, token costs, and practical limits like context window.

Sakana Fugu and Sakana Fugu-Ultra finish on the same BenchAlign overall score, so this is less about a single winner and more about where the edge shows up. The BenchAlign headline says tie; the benchmark table is where the real choice happens.

Category breakdown

Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.

Category scores and score margins for Sakana Fugu and Sakana Fugu-Ultra
CategorySakana FuguΔSakana Fugu-Ultra
ReasoningSakana Fugu86.6Margin 7.0Sakana Fugu-Ultra93.6
CodingSakana Fugu59.7Margin 4.8Sakana Fugu-Ultra64.5
AgenticSakana Fugu80.2Margin 1.9Sakana Fugu-Ultra82.1
MultimodalSakana Fugu85.1Margin 1.5Sakana Fugu-Ultra86.6
KnowledgeSakana Fugu95.5MarginTieSakana Fugu-Ultra95.5

Decisive benchmark drivers

The largest measured benchmark gaps in this matchup, with exact reported values.

More
A · Sakana FuguB · Sakana Fugu-Ultra
  1. SWE-bench Pro

    Coding
    Source ↗
    A 59%B 73.7%
    Winner: Sakana Fugu-UltraΔ 14.7
    SWE-bench Pro: Sakana Fugu scored 59%; Sakana Fugu-Ultra scored 73.7%. Sakana Fugu-Ultra wins this benchmark.
  2. MRCRv2

    Reasoning
    Source ↗
    A 86.6%B 93.6%
    Winner: Sakana Fugu-UltraΔ 7
    MRCRv2: Sakana Fugu scored 86.6%; Sakana Fugu-Ultra scored 93.6%. Sakana Fugu-Ultra wins this benchmark.
  3. Terminal-Bench 2.0

    Agentic
    Source ↗
    A 80.2%B 82.1%
    Winner: Sakana Fugu-UltraΔ 1.9
    Terminal-Bench 2.0: Sakana Fugu scored 80.2%; Sakana Fugu-Ultra scored 82.1%. Sakana Fugu-Ultra wins this benchmark.
  4. CharXiv

    Multimodal
    Source ↗
    A 85.1%B 86.6%
    Winner: Sakana Fugu-UltraΔ 1.5
    CharXiv: Sakana Fugu scored 85.1%; Sakana Fugu-Ultra scored 86.6%. Sakana Fugu-Ultra wins this benchmark.
  5. SciCode

    Coding
    Source ↗
    A 60.1%B 58.7%
    Winner: Sakana FuguΔ 1.4
    SciCode: Sakana Fugu scored 60.1%; Sakana Fugu-Ultra scored 58.7%. Sakana Fugu wins this benchmark.

Operational comparison

Runtime and commercial metrics are compared only when both models have a complete sourced value.

MetricSakana FuguSakana Fugu-UltraComparison
Input / output priceUSD per 1M tokensSakana FuguNot availableSakana Fugu-UltraNot availableA complete price comparison is not available.
Generation speedtokens per secondSakana FuguNot availableSakana Fugu-UltraNot availableA complete speed comparison is not available.
First-answer latencyseconds to first tokenSakana FuguNot availableSakana Fugu-UltraNot availableA complete latency comparison is not available.
Context windowmaximum listed tokensSakana Fugu1MSakana Fugu-Ultra1MListed context windows are equal.

Benchmark Deep Dive

AgenticSakana Fugu-Ultra wins
BenchmarkSakana FuguSakana Fugu-UltraResult
Terminal-Bench 2.0Source 80.2%82.1%Sakana Fugu-Ultra leads
CodingSakana Fugu-Ultra wins
BenchmarkSakana FuguSakana Fugu-UltraResult
SWE-bench ProSource 59%73.7%Sakana Fugu-Ultra leads
Terminal-Bench 2.0Source 80.2%82.1%Sakana Fugu-Ultra leads
LiveCodeBench v6Source 92.9%93.2%Sakana Fugu-Ultra leads
LiveCodeBench ProSource 87.8%90.8%Sakana Fugu-Ultra leads
SciCodeSource 60.1%58.7%Sakana Fugu leads
ReasoningSakana Fugu-Ultra wins
BenchmarkSakana FuguSakana Fugu-UltraResult
MRCRv2Source 86.6%93.6%Sakana Fugu-Ultra leads
KnowledgeTie
BenchmarkSakana FuguSakana Fugu-UltraResult
GPQASource 95.5%95.5%Tie
GPQA-DSource 95.5%95.5%Tie
HLE w/o toolsSource 47.2%50%Sakana Fugu-Ultra leads
MultimodalSakana Fugu-Ultra wins
BenchmarkSakana FuguSakana Fugu-UltraResult
CharXivSource 85.1%86.6%Sakana Fugu-Ultra leads
Frequently Asked Questions (6)

Which is better, Sakana Fugu or Sakana Fugu-Ultra?

Sakana Fugu and Sakana Fugu-Ultra are sibling variants in the Sakana Fugu family, so the right pick depends on whether you value the better benchmark line, cheaper tokens, or the larger context window. They are tied on the BenchAlign leaderboard on the current data.

Which is better for knowledge tasks, Sakana Fugu or Sakana Fugu-Ultra?

Sakana Fugu and Sakana Fugu-Ultra are effectively tied for knowledge tasks here, both landing at 95.5 on average.

Which is better for coding, Sakana Fugu or Sakana Fugu-Ultra?

Sakana Fugu-Ultra has the edge for coding in this comparison, averaging 64.5 versus 59.7. Inside this category, SWE-bench Pro is the benchmark that creates the most daylight between them.

Which is better for reasoning, Sakana Fugu or Sakana Fugu-Ultra?

Sakana Fugu-Ultra has the edge for reasoning in this comparison, averaging 93.6 versus 86.6. Inside this category, MRCRv2 is the benchmark that creates the most daylight between them.

Which is better for agentic tasks, Sakana Fugu or Sakana Fugu-Ultra?

Sakana Fugu-Ultra has the edge for agentic tasks in this comparison, averaging 82.1 versus 80.2. Inside this category, Terminal-Bench 2.0 is the benchmark that creates the most daylight between them.

Which is better for multimodal and grounded tasks, Sakana Fugu or Sakana Fugu-Ultra?

Sakana Fugu-Ultra has the edge for multimodal and grounded tasks in this comparison, averaging 86.6 versus 85.1. Inside this category, CharXiv is the benchmark that creates the most daylight between them.

Related Comparisons

Last updated: July 23, 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.