Sakana describes Fugu as a learned multi-model orchestration system that can route and synthesize work across underlying models. It is a product and API, not one standalone foundation model directly comparable on every task with Claude, GPT, or Gemini.13

What Fugu actually is

Sakana AI’s founders have well-documented biographies, and it is tempting to read a pedigree as a performance guarantee. Company age and employment history tell you nothing about how the model actually behaves.

Sakana positions Fugu Ultra for complex, multi-step work. Suitability for research, cybersecurity, or patent tasks requires domain-specific evaluation, data controls, tool permissions, and expert review.13

Why this, why now

The export-control narrative attached to Fugu; named rival models, a worldwide cutoff, a “route around a ban” motive; appears nowhere in Sakana’s own materials. It is a story told about the model, not by it.

TechCrunch clocked it as part of a pattern, not a one-off: Asian AI labs launching Mythos-adjacent systems while Anthropic's export ban drags on. Fugu is the most credentialed entrant so far, but it won't be the last.

Sakana publishes performance claims, but benchmark results depend on version, harness, tools, budget, and comparator settings. That makes the circulating SWE-Bench and LiveCodeBench figures unusable as a ranking until somebody reproduces them independently.13

On paper, Fugu Ultra looks formidable. It posts 73.7% on SWE-Bench Pro, ahead of Claude Opus 4.8's 69.2%, GPT-5.5's 58.6%, and Gemini 3.1 Pro's 54.2%. On LiveCodeBench it scores 93.2 against Fable 5's 89.8. It reportedly beats the older Mythos Preview on GPQA-D, the graduate-level science benchmark.

Here's the catch worth stating plainly: none of those Fable or Mythos comparisons are head-to-head. Both models were pulled from public access by the same export order that inspired Fugu's existence, so Sakana benchmarked against Anthropic's own published reference scores, not a live run. On SWE-Bench Pro specifically, Fugu Ultra actually trails Fable 5's reported 80.0; a number Sakana's own materials disclose.

AI researcher Ethan Mollick ran his usual coding tests against Fugu Ultra and found it "incredibly slow," with a single run taking 30 minutes. His verdict: results were fine, but fell short of Fable in practice.

Official API pricing lists base and higher long-context rates, while subscriptions and usage allowances can change. Orchestration work is billable, so effective task cost depends on internal calls, output length, reasoning, and retries; not only headline input/output rates.23

  • API: $5 per million input tokens, $30 per million output tokens (Fugu Ultra)
  • Standard subscription: $20/month
  • Pro subscription: $100/month, 10x usage
  • Max subscription: $200/month, 20x usage

Some estimates put Fugu Ultra's effective cost at five times Claude Opus 4.8's, for a result multiple testers describe as slower and, in practice, weaker.

The real argument

I'd argue the interesting critique isn't speed or price, it's the premise. Fugu doesn't eliminate dependency on frontier labs; it spreads that dependency across five of them and hides the seams behind one endpoint. That's a genuinely useful abstraction for a team that got burned when Fable 5 vanished from their stack overnight. It is not the same thing as building a sovereign model that doesn't need Claude, GPT, or Gemini to function, and Sakana has never claimed otherwise, whatever the headlines implied.

Fugu demonstrates that orchestration is a commercial product approach. It does not prove vendor independence, uninterrupted access, lower cost, or superior quality; underlying providers, policies, regions, and dependencies still matter.123