22 Juin

Sakana Fugu orchestrates competing models instead of building one

Sakana AI released Fugu and Fugu Ultra today, and the pitch is straightforward: instead of training yet another frontier model, build a system that coordinates them. You call one API endpoint. Behind it, a smaller language model decides which frontier models to route your request to, manages their interactions, and synthesizes outputs into a single response.

The timing is the story.

Anthropic’s Fable 5 and Mythos have been offline for ten days. The US Department of Commerce issued an emergency export control directive on June 12, barring Anthropic from distributing both models globally. Organizations that built critical infrastructure on those APIs found their access vanish overnight. Today is also the last day of Fable 5’s free trial window. Subscribers lose free access and model access simultaneously.

Enter Fugu, from a Tokyo-based lab, with the tagline: “frontier capability without the risk of export controls.”

The architecture matters here. Fugu is not a frontier model. It is a model orchestrator built on Sakana AI’s ICLR 2026 papers, Trinity and Conductor. A smaller language model sits at the core and learns when and how to call other models, including itself. It picks agents, manages their interactions, and decides which models handle which parts of a task. The agent pool is swappable. If a provider restricts access, Fugu routes around it.

Two variants launched. Fugu targets everyday work with lower latency. Fugu Ultra goes for maximum quality on hard problems.

On benchmarks, Fugu Ultra scores 73.7 on SWE-Bench Pro. That beats Opus 4.8 (69.2), GPT-5.5 (58.6), and Gemini 3.1 Pro (54.2). Sakana claims it “stands shoulder-to-shoulder” with Fable 5 and Mythos Preview. The fine print is less generous: Fable 5 scores 80.3 on the same benchmark. Fugu Ultra trails it by 6.6 points. Sakana also reports strong results on GPQA-Diamond (PhD-level scientific reasoning), LiveCodeBench, ALE-Bench, and AutoResearch.

Multiple sources note that all baseline scores for competing models are provider-reported, not independently verified in the same evaluation environment. The standard Fugu model actually outperforms Fugu Ultra on SciCode and Long Context Reasoning, suggesting the orchestration layer adds noise for document-heavy tasks.

The real claim is not about benchmarks. It is about vendor independence. Fugu’s agent pool explicitly does not include Fable 5 or Mythos. They cannot be included because they are subject to export controls. Fugu Ultra approaches their performance using only publicly available models like GPT-5.5, Gemini, and Claude Opus. If one of those gets restricted tomorrow, Fugu swaps in alternatives without rebuilding anything.

The beta period started April 25, 2026. The general availability launch today makes it Sakana AI’s first widely accessible commercial product. The API is OpenAI-compatible and available at console.sakana.ai.

Hacker News picked it up at 146 points and 89 comments within hours. Reddit’s r/singularity ran a thread. Multiple analysis pieces frame this as the start of orchestration as a product category rather than another model release.

The structural bet is simple. Modelmakers compete on raw capability. Orchestrators compete on routing intelligence. If you believe no single model will stay dominant forever, and that geopolitical access to models will keep getting more volatile, Fugu is the hedge. If you believe one model will pull ahead decisively, an orchestration layer is overhead you do not need.

Sakana AI is betting on the first scenario. The export control precedent set on June 12 makes that bet look less speculative than it did two weeks ago.

Mots-cles

sakana ai fugu orchestration export controls ai models