Routing Beats the Champion: Sakana Fugu Max and the Economics of the Open-Model Pool

Routing Beats the Champion: Sakana Fugu Max and the Economics of the Open-Model Pool

On September 11, 2026, Sakana AI released two products that say the same thing a different way: the best single model is no longer the unit that wins. Fugu Max routes every task to the cheapest open-weight model that can actually solve it, and prices the result at $2 per million input tokens and $6 per million output. Fugu Ultra v2 runs the same orchestration engine for peak quality on hard multi-step work, and reports beating Opus 5 on a visual-reasoning benchmark with none of the usual frontier names even in its pool. Neither is a model in the way the industry has measured models. Both are routers wearing a model card.

The claim that matters is not a benchmark score. It is that orchestration over many open models is now an economically serious alternative to renting one champion, and that the cost and resilience advantages come from architecture, not from a bigger training run.

The hook: two products, one engine, zero single champions

Sakana frames Max and Ultra v2 as the same core orchestration architecture aimed at two different missions. Fugu Max answers “what is the best output at the lowest cost?” Fugu Ultra v2 answers “what is the highest capability on complex, multi-step tasks?” The pool under both is a set of open-weight and specialized models, expanded by Sakana’s August 2026 collaboration with NVIDIA around its open Nemotron family. The engine, per the Sakana Fugu technical report, is itself an LLM trained to read a user query and dynamically devise an agentic scaffold, then dispatch subtasks to whatever model in the pool suits them.

That is the mechanism. A Fugu model is not a front-end prompt wrapper bolted onto one big model. It is a routing policy that assembles a working team from the pool per query, which is precisely the “models split by job” idea applied at the product level instead of at the research level.

The economics: 40 to 60 percent cheaper output, routed down

The most concrete number in the release is cost. Fugu Max prices at $2 per million input and $6 per million output tokens. Sakana states its output pricing is 40 to 60 percent below Sonnet 5, GPT 5.6 Terra, and Kimi K3. That is a routing win, not a training win: the pool stores many small, specialized models, and the engine sends each subproblem to the leanest member that can handle it. A task that the frontier would process with a trillion-plus parameter model gets split into pieces priced like their actual difficulty.

Sakana runs the same cost argument on benchmarks. Fugu Max reports the best overall score on six evaluations: Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and SWEFish. It says it expands the cost-performance Pareto frontier on seven of ten benchmarks. The framing is deliberate: the frontier the industry should care about has two axes, capability and cost, and single-model providers can only move along one by picking a bigger champion.

The peak flavor: beating Opus 5 without the usual suspects

Fugu Ultra v2 is the capability story, and its most striking detail is how little it leans on the closed frontier. It reports 48.3 on Chartography, a visual-reasoning and data-interpretation benchmark, against Opus 5’s 27.3 and Fable 5’s 29.5. On DeepSWE, a real-world software-engineering benchmark, it reports 74.3 against models that cost three to five times more per token. Sakana is explicit that these scores come with no Fable 5, Fable 5.1, or GPT-6-Astra in the agent pool, and that the pool is swappable by design.

Read that line as a map, not just a boast. It tells you Sakana currently cannot draw on those particular frontier stacks, or chooses not to, and still reports top results. That is the architectural point under the marketing: capability can come from composing open models, and the closed frontier, however strong, is no longer the only route to a state-of-the-art score. The supply chain is the differentiator. An open, swappable pool survives an API revocation, a model cutoff, or a licensing change, where a system wired to one champion does not. Sakana calls this “supply chain resilience by design” and frames it as the infrastructure for AI sovereignty.

What common intuition gets wrong

The default intuition is that orchestration, routing, or a model pool is a cost optimization you apply after you pick the strongest model. Fugu Max treats routing as the product, and treats the model pool as the asset you maintain. The uncomfortable consequence for operators: a router loses most of its value if it only decides which vendor’s frontier model to call, because routing between two champions still pays champion prices. The value comes from a pool with genuinely cheap, small, specialized members, and from the discipline to actually send easy tasks to them.

A second wrong intuition is to treat self-reported orchestration scores as comparable to a single-model leaderboard run. Every Fugu Max and Ultra v2 number in the announcement is Sakana’s own; no independent third-party evaluation is cited. On agentic benchmarks especially, methodology matters as much as the engine, and a vendor-controlled harness around a routing pool is a different measurement surface than a fixed question set. The right reading is directional: orchestration over open models is now competitive and dramatically cheaper, not that it definitively beats the frontier on neutral ground.

The operator consequence

For anyone building agents or buying model capacity, the change is in the default questions you should ask:

  1. Stop asking which model is best. Ask which pool and router are best. If a routing layer over open and specialized models delivers Frontier-adjacent quality at 40 to 60 percent lower output cost, then your single-vendor integration is a fallback, not a default.

  2. Design for a swappable pool. Pin your agent stack to concrete model IDs and a single closed API and you inherit that vendor’s revocations and price changes. A pool you control, with routing as a policy you can tune, converts supplier risk into a config value.

  3. Buy the router’s economics, verify the router’s numbers. Treat cost claims as direction, validate quality against your own workload, and instrument the agentic task distribution. The real deliverable is not a score; it is that easy subtasks stop paying champion prices.

  4. Watch the boundary, because orchestration enlarges it. A router that dispatches across many models has many trust surfaces: several model providers, more tool schemas, more intermediate scaffolds. Pool economics are only worth it if you can see and sandbox every member.

Closing thesis

The capability frontier has not vanished. It has moved. It is now a property of the routing and orchestration layer that assembles a team of open models, not of any single set of weights. Sakana Fugu Max is the cleanest recent proof that composing cheap, specialized models under a smart router beats renting a single champion, on cost, on resilience, and increasingly on raw score. The model you picked was never really the whole problem. The pool, the router, and the supply chain around them are the system now, and they are open.

Sources:

Keep reading