Last updated: July 24, 2026
Sakana AI shipped Fugu-Ultra v1.1 on July 24, 2026, a little over a month after the original Fugu and Fugu Ultra launch on June 22. The pitch in Sakana’s official announcement on X: newer frontier models inside Fugu’s agent pool, stronger numbers across every benchmark Sakana shows, and the same price as v1.0.
Quick answer: Fugu-Ultra v1.1 is a drop-in refresh of Sakana’s quality-first orchestration model, not a new product line. Sakana says v1.1 gains up to 7.9 points over v1.0 across its published benchmark suite, with the strongest improvements on ProgramBench and Terminal Bench 2.1, and that it is “more capable across coding, agentic tasks, and advanced reasoning.” Pricing on the Sakana Fugu product page is listed jointly for fugu-ultra-v1.1 and fugu-ultra-v1.0 (the renamed fugu-ultra-20260615): $5 per 1M input tokens, $30 per 1M output tokens, $0.50 cached input. The catch: the v1.1 numbers are vendor-published, and Sakana has not released a per-benchmark v1.0-vs-v1.1 changelog at the time of writing.
| Fugu-Ultra v1.1 fact | Detail |
|---|---|
| Release date | July 24, 2026, per Sakana’s X announcement |
| What changed | Newer frontier models in Fugu’s agent pool; Sakana claims gains on every benchmark it shows |
| Headline gain | Up to +7.9 points over v1.0 |
| Strongest improvements | ProgramBench and Terminal Bench 2.1, per Sakana |
| Pricing | Unchanged: $5 input / $30 output / $0.50 cached per 1M tokens; double above 272K context |
| Model IDs | fugu-ultra-v1.1 (new) and fugu-ultra-v1.0 (renamed fugu-ultra-20260615) |
| Family status | Fugu, Fugu Ultra, and Fugu Cyber all listed on sakana.ai/fugu |
| Access | OpenAI-compatible API, plus OpenRouter, Vercel AI Gateway, and models.dev listings |
| Not available | EU/EEA while Sakana works on GDPR compliance |
| Evidence level | Vendor-published claims; no independent v1.1 benchmarks yet |
Source check — July 24, 2026: this article checks Sakana’s official Fugu-Ultra v1.1 announcement on X, the Sakana Fugu product page for pricing, model IDs, family lineup, and the published benchmark table, the original Fugu launch post for v1.0 context, the Fugu technical report, and the SakanaAI/fugu GitHub repository. Because orchestration products can swap underlying workers at any time, verify the live Sakana Fugu pricing page and console before changing production traffic.
For the broader context on how Fugu works as a multi-agent orchestration system, start with the Sakana AI Fugu review. For the standing Fugu Ultra guide, see Fugu Ultra: model, pricing, benchmarks, and use cases. For coding-tool setup, see Sakana Fugu with Codex and Cursor.
What’s New in Fugu-Ultra v1.1
Sakana’s announcement is short, and the substance is in three claims:
- Newer frontier models in the pool. Sakana says v1.1 is “upgraded to incorporate the latest frontier models.” Because Fugu is an orchestration model — a language model trained to route work across a pool of specialist agents — refreshing the pool is the main lever for improving quality without retraining the orchestrator from scratch.
- Gains across every benchmark Sakana shows. Sakana claims “stronger performance across every benchmark shown, including gains of up to 7.9 points over v1.0.” The “up to” phrasing matters: 7.9 is the maximum delta Sakana reports, not the average.
- Coding and agentic tasks move the most. Sakana singles out ProgramBench and Terminal Bench 2.1 as particularly strong improvement areas, and describes v1.1 as “more capable across coding, agentic tasks, and advanced reasoning.”
What Sakana did not publish at launch:
- a per-benchmark v1.0-vs-v1.1 score table;
- which specific frontier models entered or left the agent pool;
- whether the orchestrator itself was retrained or only the pool refreshed;
- independent third-party v1.1 results.
That is normal for a point release, but it changes how the claims should be read. “Up to 7.9 points” is a ceiling, and ProgramBench and Terminal Bench 2.1 are the categories Sakana chose to highlight. Until Sakana publishes a full delta table — or independent benchmarkers run v1.1 — treat the gain distribution as marketing-shaped.
The Quiet Change: fugu-ultra-20260615 Is Now fugu-ultra-v1.0
The most operationally important detail is on the Sakana Fugu product page, not in the tweet. Pricing is now listed as:
Fixed pricing for fugu-ultra-v1.1 and fugu-ultra-v1.0 (previously fugu-ultra-20260615)
That formalizes a model-ID rename. The original June 2026 Fugu Ultra endpoint, which our standing Fugu Ultra guide documented as fugu-ultra-20260615, is now fugu-ultra-v1.0. Teams that hardcoded the dated ID should check whether their integration still resolves it and migrate to the versioned IDs.
| Model ID | Status as of July 24, 2026 |
|---|---|
fugu-ultra-v1.1 | Current quality-first Fugu Ultra model |
fugu-ultra-v1.0 | Previous Fugu Ultra model; formerly fugu-ultra-20260615 |
fugu-ultra-20260615 | Old dated ID; Sakana’s docs now refer to this as v1.0 |
The versioned naming is a small but real signal of product maturity: Sakana is moving from dated snapshots to a v1.x series, which makes drop-in upgrades and rollbacks easier to reason about.
Fugu-Ultra v1.1 Benchmarks: What Sakana Publishes
The Sakana Fugu product page currently shows the following Fugu Ultra benchmark table. Sakana does not explicitly label the table as v1.0 or v1.1 at the time of writing, so read it as Sakana’s current published Fugu Ultra picture, with v1.1 claimed to improve on it:
| Benchmark | Fugu (standard) | Fugu Ultra | Opus 4.8 † | Gemini 3.1 Pro † | GPT 5.5 † |
|---|---|---|---|---|---|
| SWE Bench Pro * | 59.0 | 73.7 | 69.2 | 54.2 | 58.6 |
| TerminalBench 2.1 | 80.2 | 82.1 | 74.6 | 70.3 | 78.2 |
| LiveCodeBench | 92.9 | 93.2 | 87.8 | 88.5 | 85.3 |
| LiveCodeBench Pro | 87.8 | 90.8 | 84.8 | 82.9 | 88.4 |
| Humanity’s Last Exam | 47.2 | 50.0 | 49.8 | 44.4 | 41.4 |
| CharXiv Reasoning | 85.1 | 86.6 | 84.2 | 83.3 | 84.1 |
| GPQA-D | 95.5 | 95.5 | 92.0 | 94.3 | 93.6 |
| SciCode | 60.1 | 58.7 | 53.5 | 58.9 | 56.1 |
| τ³ Banking | 21.7 | 20.6 | 20.6 | 8.4 | 20.6 |
| Long Context Reasoning | 74.7 | 73.3 | 67.7 | 72.7 | 74.3 |
| MRCRv2 | 86.6 | 93.6 | 87.9 | 84.9 | 94.8 |
* Sakana uses mini-swe-agent as the scaffolding for SWE Bench Pro. † Baseline scores are model provider-reported, per Sakana’s table notes.
Two honest caveats from Sakana’s own framing:
- Baselines are provider-reported. Sakana does not re-run Opus 4.8, Gemini 3.1 Pro, or GPT 5.5; it cites their publishers’ numbers.
- Fable 5 and Mythos Preview are not in Fugu’s pool because they are not publicly accessible; Sakana reports the max of the two where both have a score on the same benchmark.
The v1.1 announcement adds two specific claims on top of this table: gains of up to 7.9 points over v1.0, and particularly strong movement on ProgramBench (not in the table above) and Terminal Bench 2.1 (82.1 for Fugu Ultra in the current table, versus 78.2 for GPT 5.5 and 74.6 for Opus 4.8). If Sakana’s v1.1 Terminal Bench gain is measured against v1.0’s own score, the v1.1 number should be meaningfully above 82.1 — but Sakana has not published the exact v1.1 figure at the time of writing.
| Signal | What it supports | What it does not prove |
|---|---|---|
| “Up to 7.9 points over v1.0” | Real ceiling improvement on at least one benchmark | Average improvement across the suite |
| ProgramBench and Terminal Bench 2.1 highlighted | Coding and terminal/agentic tasks likely improved most | That every coding task improves uniformly |
| Same price as v1.0 | Cost-per-task should not rise from the model swap alone | That orchestration token usage is unchanged |
| Vendor-published table | Sakana’s own evaluation methodology | Independent reproducibility |
Pricing: Same Money, Newer Model
The cleanest part of the v1.1 story is that pricing is unchanged. Sakana lists one fixed price card covering both fugu-ultra-v1.1 and fugu-ultra-v1.0:
| Token type | Price per 1M tokens | Price when context is over 272K |
|---|---|---|
| Input | $5 | $10 |
| Output | $30 | $45 |
| Cached input | $0.50 | $1.00 |
Subscription tiers also remain the same, and every tier includes both standard Fugu and Fugu Ultra:
| Plan | Price | Positioning |
|---|---|---|
| Standard | $20/month | Lightweight daily usage |
| Pro | $100/month | Focused coding, review, research, and analysis sessions |
| Max | $200/month | Heavier long-running workloads |
Sakana is also running a launch-window promotion: subscribe before the end of July 2026 and the second month at your initial tier is free, per the product page.
The number that actually matters for production is still cost per completed task, not sticker price. Fugu Ultra orchestrates multiple agents behind one API call, and Sakana’s pricing notes explain that you are charged a single rate based on the top-tier model involved rather than stacked per-model fees. That is better than naive multi-model billing, but orchestration can still expand total token usage. If you re-test v1.1, log input, output, cached-input, latency, retry rate, and human review time alongside the headline answer quality.
For side-by-side pricing context against the premium single-model alternatives, compare with Claude Fable 5 and Claude Fable 5 vs Opus 4.8.
Fugu Cyber: The Family Now Has Three Models
The other change visible on the Fugu product page since the June launch is the family shape. Sakana now lists three Fugu models behind the same OpenAI-compatible API:
| Model | Positioning | Token Plan pricing (per 1M tokens) |
|---|---|---|
| Fugu | Balanced performance and latency; everyday coding, code review, chatbots | Standard rate of the underlying top-tier model; no stacked fees |
| Fugu Ultra | Maximum answer quality on hard multi-step work; research, paper reproduction, patent and literature analysis | $5 input / $30 output / $0.50 cached (v1.1 and v1.0 share the card) |
| Fugu Cyber | Cybersecurity reasoning: vulnerability research, threat investigation, security analysis | $6 input / $36 output / $0.60 cached; Token Plan only |
Sakana says Fugu Cyber reaches 86.9% on CyberGym and 72.1% on CTI-REALM, which it describes as comparable to GPT-5.5-Cyber and Mythos-Preview. As with the rest of Sakana’s numbers, these are vendor-published.
The practical implication: Fugu-Ultra v1.1 is not the only thing that changed in the family since our June coverage. Teams doing scoped security review should look at Fugu Cyber as a separate evaluation track rather than assuming Ultra covers it.
Availability: API, OpenRouter, Vercel — but Not the EU
Sakana Fugu, including Fugu Ultra, is accessible through:
- Sakana’s own OpenAI-compatible API via the console — point an existing OpenAI client or coding harness at the Fugu endpoint, no SDK migration required;
- OpenRouter;
- Vercel’s AI Gateway;
- models.dev provider listings.
The standing restriction has not changed: Sakana Fugu is not yet available in the EU/EEA while Sakana works toward GDPR and EU-specific compliance, per the product page. For teams with EU users or EU data-processing requirements, that is still a hard gate, regardless of how good the v1.1 numbers look.
For coding-tool wiring (Codex, Cursor, Claude Code caveats), see Sakana Fugu with Codex and Cursor.
Should You Move to v1.1?
Because pricing is identical and the API is OpenAI-compatible, the upgrade question is mostly about evaluation effort, not procurement.
- You already run Fugu Ultra and can A/B model IDs
- Your workload is coding, terminal, or agentic — the areas Sakana highlights
- You have a reusable internal task suite to re-score
- You only have anecdotal prompts to test with
- Your tasks are long-context (>272K) where pricing doubles
- You depend on the old fugu-ultra-20260615 ID and have not verified aliasing
- You need independent v1.1 benchmarks before adopting
- You serve EU/EEA users
- You require per-request visibility into which underlying models were used
A practical re-test checklist:
| Check | What to log |
|---|---|
| Answer quality | Human-scored pass rate on 20-50 representative tasks, v1.0 vs v1.1 |
| Coding/terminal subset | Separate score for the categories Sakana claims improved most |
| Cost per task | Input, output, cached tokens, and any orchestration usage |
| Latency | P50 and P95 end-to-end response time |
| Retry behavior | How often v1.1 needs a second prompt versus v1.0 |
| Failure modes | Refusals, scope drift, broken code, hallucinated citations |
If v1.1 wins or ties on your suite, switching costs nothing extra. If it regresses on a narrow category, Sakana still lists v1.0 on the same price card, so pinning is an option rather than a dead end.
How v1.1 Fits the Frontier Picture
Fugu-Ultra v1.1 lands in a crowded month for frontier and near-frontier models. The structural argument for Fugu has not changed since the June launch: instead of betting on one provider’s model, you bet on an orchestrator that can swap workers as the frontier moves. The v1.1 release is the first concrete demonstration of that thesis playing out in production — Sakana refreshed the pool and passed the gains through at the same price.
The counter-arguments have not changed either:
- Vendor-published benchmarks are not independent proof. The same caveat from our Sakana AI Fugu review applies to v1.1: Sakana’s table is evidence, not verification.
- Routing opacity is a governance cost. You still cannot see exactly which underlying models handled a given request, which matters for compliance-heavy environments.
- Single-model alternatives keep shipping. Claude Opus 4.8, Claude Fable 5, and the GPT-5.5/5.6 line are all moving targets, and Kimi K3 just pressured the open-weight side.
For developers choosing a coding agent rather than a model API, the relevant comparison is less “v1.1 vs Fable 5” and more “orchestrated API vs your harness of choice” — see Claude Code vs Codex and best AI agents for coding for that layer.
Final Verdict: A Free Quality Reroll, With the Usual Vendor-Claim Asterisk
Fugu-Ultra v1.1 is exactly what a point release should be: newer workers, claimed gains on the benchmarks that matter for coding and agentic work, unchanged pricing, and a clean versioned model ID. For existing Fugu Ultra users, the upgrade is low-risk because the price card is shared and the old v1.0 remains listed. For new evaluators, v1.1 makes Fugu Ultra a slightly stronger candidate against Fable 5, Opus 4.8, and GPT-5.5/5.6 — but the evidence is still Sakana’s own.
The right posture is the same as at the June launch, only more specific: test v1.1 on your own task suite, especially coding and terminal-style work, before believing the +7.9 headline. If Sakana publishes a full v1.0-vs-v1.1 delta table, or independent benchmarkers reproduce the gains, the story gets stronger. Until then, v1.1 is a promising refresh, not a verified frontier shift.