Last updated: July 24, 2026

Sakana AI shipped Fugu-Ultra v1.1 on July 24, 2026, a little over a month after the original Fugu and Fugu Ultra launch on June 22. The pitch in Sakana’s official announcement on X: newer frontier models inside Fugu’s agent pool, stronger numbers across every benchmark Sakana shows, and the same price as v1.0.

Quick answer: Fugu-Ultra v1.1 is a drop-in refresh of Sakana’s quality-first orchestration model, not a new product line. Sakana says v1.1 gains up to 7.9 points over v1.0 across its published benchmark suite, with the strongest improvements on ProgramBench and Terminal Bench 2.1, and that it is “more capable across coding, agentic tasks, and advanced reasoning.” Pricing on the Sakana Fugu product page is listed jointly for fugu-ultra-v1.1 and fugu-ultra-v1.0 (the renamed fugu-ultra-20260615): $5 per 1M input tokens, $30 per 1M output tokens, $0.50 cached input. The catch: the v1.1 numbers are vendor-published, and Sakana has not released a per-benchmark v1.0-vs-v1.1 changelog at the time of writing.

Fugu-Ultra v1.1 factDetail
Release dateJuly 24, 2026, per Sakana’s X announcement
What changedNewer frontier models in Fugu’s agent pool; Sakana claims gains on every benchmark it shows
Headline gainUp to +7.9 points over v1.0
Strongest improvementsProgramBench and Terminal Bench 2.1, per Sakana
PricingUnchanged: $5 input / $30 output / $0.50 cached per 1M tokens; double above 272K context
Model IDsfugu-ultra-v1.1 (new) and fugu-ultra-v1.0 (renamed fugu-ultra-20260615)
Family statusFugu, Fugu Ultra, and Fugu Cyber all listed on sakana.ai/fugu
AccessOpenAI-compatible API, plus OpenRouter, Vercel AI Gateway, and models.dev listings
Not availableEU/EEA while Sakana works on GDPR compliance
Evidence levelVendor-published claims; no independent v1.1 benchmarks yet

Source check — July 24, 2026: this article checks Sakana’s official Fugu-Ultra v1.1 announcement on X, the Sakana Fugu product page for pricing, model IDs, family lineup, and the published benchmark table, the original Fugu launch post for v1.0 context, the Fugu technical report, and the SakanaAI/fugu GitHub repository. Because orchestration products can swap underlying workers at any time, verify the live Sakana Fugu pricing page and console before changing production traffic.

For the broader context on how Fugu works as a multi-agent orchestration system, start with the Sakana AI Fugu review. For the standing Fugu Ultra guide, see Fugu Ultra: model, pricing, benchmarks, and use cases. For coding-tool setup, see Sakana Fugu with Codex and Cursor.

What’s New in Fugu-Ultra v1.1

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Sakana’s announcement is short, and the substance is in three claims:

  1. Newer frontier models in the pool. Sakana says v1.1 is “upgraded to incorporate the latest frontier models.” Because Fugu is an orchestration model — a language model trained to route work across a pool of specialist agents — refreshing the pool is the main lever for improving quality without retraining the orchestrator from scratch.
  2. Gains across every benchmark Sakana shows. Sakana claims “stronger performance across every benchmark shown, including gains of up to 7.9 points over v1.0.” The “up to” phrasing matters: 7.9 is the maximum delta Sakana reports, not the average.
  3. Coding and agentic tasks move the most. Sakana singles out ProgramBench and Terminal Bench 2.1 as particularly strong improvement areas, and describes v1.1 as “more capable across coding, agentic tasks, and advanced reasoning.”

What Sakana did not publish at launch:

  • a per-benchmark v1.0-vs-v1.1 score table;
  • which specific frontier models entered or left the agent pool;
  • whether the orchestrator itself was retrained or only the pool refreshed;
  • independent third-party v1.1 results.

That is normal for a point release, but it changes how the claims should be read. “Up to 7.9 points” is a ceiling, and ProgramBench and Terminal Bench 2.1 are the categories Sakana chose to highlight. Until Sakana publishes a full delta table — or independent benchmarkers run v1.1 — treat the gain distribution as marketing-shaped.

The Quiet Change: fugu-ultra-20260615 Is Now fugu-ultra-v1.0

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

The most operationally important detail is on the Sakana Fugu product page, not in the tweet. Pricing is now listed as:

Fixed pricing for fugu-ultra-v1.1 and fugu-ultra-v1.0 (previously fugu-ultra-20260615)

That formalizes a model-ID rename. The original June 2026 Fugu Ultra endpoint, which our standing Fugu Ultra guide documented as fugu-ultra-20260615, is now fugu-ultra-v1.0. Teams that hardcoded the dated ID should check whether their integration still resolves it and migrate to the versioned IDs.

Model IDStatus as of July 24, 2026
fugu-ultra-v1.1Current quality-first Fugu Ultra model
fugu-ultra-v1.0Previous Fugu Ultra model; formerly fugu-ultra-20260615
fugu-ultra-20260615Old dated ID; Sakana’s docs now refer to this as v1.0

The versioned naming is a small but real signal of product maturity: Sakana is moving from dated snapshots to a v1.x series, which makes drop-in upgrades and rollbacks easier to reason about.

Fugu-Ultra v1.1 Benchmarks: What Sakana Publishes

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

The Sakana Fugu product page currently shows the following Fugu Ultra benchmark table. Sakana does not explicitly label the table as v1.0 or v1.1 at the time of writing, so read it as Sakana’s current published Fugu Ultra picture, with v1.1 claimed to improve on it:

BenchmarkFugu (standard)Fugu UltraOpus 4.8 †Gemini 3.1 Pro †GPT 5.5 †
SWE Bench Pro *59.073.769.254.258.6
TerminalBench 2.180.282.174.670.378.2
LiveCodeBench92.993.287.888.585.3
LiveCodeBench Pro87.890.884.882.988.4
Humanity’s Last Exam47.250.049.844.441.4
CharXiv Reasoning85.186.684.283.384.1
GPQA-D95.595.592.094.393.6
SciCode60.158.753.558.956.1
τ³ Banking21.720.620.68.420.6
Long Context Reasoning74.773.367.772.774.3
MRCRv286.693.687.984.994.8

* Sakana uses mini-swe-agent as the scaffolding for SWE Bench Pro. † Baseline scores are model provider-reported, per Sakana’s table notes.

Two honest caveats from Sakana’s own framing:

  • Baselines are provider-reported. Sakana does not re-run Opus 4.8, Gemini 3.1 Pro, or GPT 5.5; it cites their publishers’ numbers.
  • Fable 5 and Mythos Preview are not in Fugu’s pool because they are not publicly accessible; Sakana reports the max of the two where both have a score on the same benchmark.

The v1.1 announcement adds two specific claims on top of this table: gains of up to 7.9 points over v1.0, and particularly strong movement on ProgramBench (not in the table above) and Terminal Bench 2.1 (82.1 for Fugu Ultra in the current table, versus 78.2 for GPT 5.5 and 74.6 for Opus 4.8). If Sakana’s v1.1 Terminal Bench gain is measured against v1.0’s own score, the v1.1 number should be meaningfully above 82.1 — but Sakana has not published the exact v1.1 figure at the time of writing.

Evaluation template How to read the v1.1 benchmark claims
SignalWhat it supportsWhat it does not prove
“Up to 7.9 points over v1.0”Real ceiling improvement on at least one benchmarkAverage improvement across the suite
ProgramBench and Terminal Bench 2.1 highlightedCoding and terminal/agentic tasks likely improved mostThat every coding task improves uniformly
Same price as v1.0Cost-per-task should not rise from the model swap aloneThat orchestration token usage is unchanged
Vendor-published tableSakana’s own evaluation methodologyIndependent reproducibility

Pricing: Same Money, Newer Model

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

The cleanest part of the v1.1 story is that pricing is unchanged. Sakana lists one fixed price card covering both fugu-ultra-v1.1 and fugu-ultra-v1.0:

Token typePrice per 1M tokensPrice when context is over 272K
Input$5$10
Output$30$45
Cached input$0.50$1.00

Subscription tiers also remain the same, and every tier includes both standard Fugu and Fugu Ultra:

PlanPricePositioning
Standard$20/monthLightweight daily usage
Pro$100/monthFocused coding, review, research, and analysis sessions
Max$200/monthHeavier long-running workloads

Sakana is also running a launch-window promotion: subscribe before the end of July 2026 and the second month at your initial tier is free, per the product page.

The number that actually matters for production is still cost per completed task, not sticker price. Fugu Ultra orchestrates multiple agents behind one API call, and Sakana’s pricing notes explain that you are charged a single rate based on the top-tier model involved rather than stacked per-model fees. That is better than naive multi-model billing, but orchestration can still expand total token usage. If you re-test v1.1, log input, output, cached-input, latency, retry rate, and human review time alongside the headline answer quality.

For side-by-side pricing context against the premium single-model alternatives, compare with Claude Fable 5 and Claude Fable 5 vs Opus 4.8.

Fugu Cyber: The Family Now Has Three Models

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

The other change visible on the Fugu product page since the June launch is the family shape. Sakana now lists three Fugu models behind the same OpenAI-compatible API:

ModelPositioningToken Plan pricing (per 1M tokens)
FuguBalanced performance and latency; everyday coding, code review, chatbotsStandard rate of the underlying top-tier model; no stacked fees
Fugu UltraMaximum answer quality on hard multi-step work; research, paper reproduction, patent and literature analysis$5 input / $30 output / $0.50 cached (v1.1 and v1.0 share the card)
Fugu CyberCybersecurity reasoning: vulnerability research, threat investigation, security analysis$6 input / $36 output / $0.60 cached; Token Plan only

Sakana says Fugu Cyber reaches 86.9% on CyberGym and 72.1% on CTI-REALM, which it describes as comparable to GPT-5.5-Cyber and Mythos-Preview. As with the rest of Sakana’s numbers, these are vendor-published.

The practical implication: Fugu-Ultra v1.1 is not the only thing that changed in the family since our June coverage. Teams doing scoped security review should look at Fugu Cyber as a separate evaluation track rather than assuming Ultra covers it.

Availability: API, OpenRouter, Vercel — but Not the EU

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Sakana Fugu, including Fugu Ultra, is accessible through:

  • Sakana’s own OpenAI-compatible API via the console — point an existing OpenAI client or coding harness at the Fugu endpoint, no SDK migration required;
  • OpenRouter;
  • Vercel’s AI Gateway;
  • models.dev provider listings.

The standing restriction has not changed: Sakana Fugu is not yet available in the EU/EEA while Sakana works toward GDPR and EU-specific compliance, per the product page. For teams with EU users or EU data-processing requirements, that is still a hard gate, regardless of how good the v1.1 numbers look.

For coding-tool wiring (Codex, Cursor, Claude Code caveats), see Sakana Fugu with Codex and Cursor.

Should You Move to v1.1?

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Because pricing is identical and the API is OpenAI-compatible, the upgrade question is mostly about evaluation effort, not procurement.

Decision point Should you switch to Fugu-Ultra v1.1?
Best fit
  • You already run Fugu Ultra and can A/B model IDs
  • Your workload is coding, terminal, or agentic — the areas Sakana highlights
  • You have a reusable internal task suite to re-score
Use carefully
  • You only have anecdotal prompts to test with
  • Your tasks are long-context (>272K) where pricing doubles
  • You depend on the old fugu-ultra-20260615 ID and have not verified aliasing
Use another option
  • You need independent v1.1 benchmarks before adopting
  • You serve EU/EEA users
  • You require per-request visibility into which underlying models were used
Treat v1.1 as a free reroll on quality at the same price, but verify it on your own tasks. A same-price point release with vendor-claimed gains is exactly the case where a 20-50 task internal benchmark pays for itself.

A practical re-test checklist:

CheckWhat to log
Answer qualityHuman-scored pass rate on 20-50 representative tasks, v1.0 vs v1.1
Coding/terminal subsetSeparate score for the categories Sakana claims improved most
Cost per taskInput, output, cached tokens, and any orchestration usage
LatencyP50 and P95 end-to-end response time
Retry behaviorHow often v1.1 needs a second prompt versus v1.0
Failure modesRefusals, scope drift, broken code, hallucinated citations

If v1.1 wins or ties on your suite, switching costs nothing extra. If it regresses on a narrow category, Sakana still lists v1.0 on the same price card, so pinning is an option rather than a dead end.

How v1.1 Fits the Frontier Picture

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Fugu-Ultra v1.1 lands in a crowded month for frontier and near-frontier models. The structural argument for Fugu has not changed since the June launch: instead of betting on one provider’s model, you bet on an orchestrator that can swap workers as the frontier moves. The v1.1 release is the first concrete demonstration of that thesis playing out in production — Sakana refreshed the pool and passed the gains through at the same price.

The counter-arguments have not changed either:

  • Vendor-published benchmarks are not independent proof. The same caveat from our Sakana AI Fugu review applies to v1.1: Sakana’s table is evidence, not verification.
  • Routing opacity is a governance cost. You still cannot see exactly which underlying models handled a given request, which matters for compliance-heavy environments.
  • Single-model alternatives keep shipping. Claude Opus 4.8, Claude Fable 5, and the GPT-5.5/5.6 line are all moving targets, and Kimi K3 just pressured the open-weight side.

For developers choosing a coding agent rather than a model API, the relevant comparison is less “v1.1 vs Fable 5” and more “orchestrated API vs your harness of choice” — see Claude Code vs Codex and best AI agents for coding for that layer.

Final Verdict: A Free Quality Reroll, With the Usual Vendor-Claim Asterisk

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.

Fugu-Ultra v1.1 is exactly what a point release should be: newer workers, claimed gains on the benchmarks that matter for coding and agentic work, unchanged pricing, and a clean versioned model ID. For existing Fugu Ultra users, the upgrade is low-risk because the price card is shared and the old v1.0 remains listed. For new evaluators, v1.1 makes Fugu Ultra a slightly stronger candidate against Fable 5, Opus 4.8, and GPT-5.5/5.6 — but the evidence is still Sakana’s own.

The right posture is the same as at the June launch, only more specific: test v1.1 on your own task suite, especially coding and terminal-style work, before believing the +7.9 headline. If Sakana publishes a full v1.0-vs-v1.1 delta table, or independent benchmarkers reproduce the gains, the story gets stronger. Until then, v1.1 is a promising refresh, not a verified frontier shift.

FAQ

Try it in practice Make this section actionable Practice the workflow instead of only comparing tools.
What is Fugu-Ultra v1.1?
Fugu-Ultra v1.1 is the July 24, 2026 update to Sakana AI’s quality-first Fugu Ultra model. Sakana says it incorporates the latest frontier models in Fugu’s agent pool and improves across coding, agentic tasks, and advanced reasoning, with gains of up to 7.9 benchmark points over v1.0.
How much better is Fugu-Ultra v1.1 than v1.0?
Sakana claims gains of up to 7.9 points over v1.0 across the benchmarks it shows, with particularly strong results on ProgramBench and Terminal Bench 2.1. “Up to” is the maximum reported gain, not the average, and Sakana has not published a per-benchmark v1.0-vs-v1.1 table at the time of writing.
Did Fugu-Ultra pricing change with v1.1?
No. Sakana lists one fixed price card for both fugu-ultra-v1.1 and fugu-ultra-v1.0: $5 per 1M input tokens, $30 per 1M output tokens, and $0.50 per 1M cached input tokens, doubling to $10 / $45 / $1.00 when context exceeds 272K tokens.
What happened to the fugu-ultra-20260615 model ID?
Sakana’s product page now refers to the original June 2026 Fugu Ultra model as fugu-ultra-v1.0, previously fugu-ultra-20260615. If your integration hardcodes the dated ID, verify that it still resolves and migrate to the versioned IDs.
Are the v1.1 benchmark numbers independently verified?
Not at the time of writing. The v1.1 claims come from Sakana’s own announcement and product page, and baseline scores in Sakana’s table are provider-reported. Independent third-party v1.1 results have not been published yet.
Is Fugu-Ultra v1.1 available in the EU?
No. Sakana says Fugu is not yet available in the EU/EEA while it works toward GDPR and EU-specific regulatory compliance. That restriction applies to v1.1 the same way it applied at the June launch.
What is Fugu Cyber?
Fugu Cyber is a third Fugu family model specialized for cybersecurity reasoning — vulnerability research, threat investigation, and security analysis. Sakana lists it at $6 input / $36 output / $0.60 cached per 1M tokens, available only on the Token Plan, and claims 86.9% on CyberGym and 72.1% on CTI-REALM.
Should I switch from v1.0 to v1.1?
If you already run Fugu Ultra, an A/B test of the two model IDs is low-risk because pricing is identical and both remain listed. Re-run 20-50 representative tasks, with a separate slice for coding and terminal-style work, and compare answer quality, cost per completed task, latency, and retry rate before moving all traffic.