Frontier Models
Sakana AI Launches Fugu Max and Fugu Ultra v2 to Advance Multi-Agent Orchestration
The September 11, 2026 releases demonstrate that orchestration of open models can deliver top benchmark results at lower costs while avoiding dependence on closed frontier systems such as GPT-6 Astra and Claude Fable 5.1.
Fugu Max and Fugu Ultra v2 are multi-agent orchestration systems from Sakana AI that separately target cost efficiency and maximum capability.
Sakana AI made the announcement of Fugu Max and Fugu Ultra v2 on September 11, 2026. The two systems represent different points on the cost-performance spectrum. Fugu Max prioritizes efficiency while Fugu Ultra v2 prioritizes capability. This strategy allows users to select based on their specific needs for either budget-conscious applications or those requiring maximum accuracy on difficult tasks. The release expands the options available in the frontier models space by showing results from open model pools.
What background and context surround the Fugu Max and Fugu Ultra v2 releases?
Multi-agent orchestration has gained attention as a way to combine multiple models for complex tasks. Sakana AI builds on this by showing results that match or exceed those from monolithic closed models. The approach relies on a pool of models that does not include certain closed systems. Prior developments in agent frameworks often defaulted to the largest proprietary models for peak accuracy on software and tool-use evaluations.
The release statement from Sakana AI notes the simultaneous push along cost and performance axes. This comes as the industry sees increasing use of agents for software engineering and tool use benchmarks. Fugu systems demonstrate viability without the highest-end closed models. The strategy addresses both enterprise cost pressures and research demands for high capability on multi-step problems.
What new capabilities do Fugu Max and Fugu Ultra v2 bring in detail?
Fugu Ultra v2 records top or joint-top scores on five of eight benchmarks. These benchmarks include DeepSWE and Toolathon. The model handles tasks with a 1M context window and produces up to 128K tokens in output. The system processes text and image inputs to support visual reasoning alongside language tasks.
Fugu Max secures the best overall score on six benchmarks. The list includes Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and SWEFish. Its pricing structure supports broader adoption through lower rates. The dual offering allows organizations to match system choice to workload requirements.
Both systems operate through an OpenAI-compatible API. Users can switch to the new models with a single parameter adjustment. This design lowers the barrier for integration into existing workflows. Availability extends through platforms that aggregate multiple model providers.
What technical specifics characterize the Fugu models?
Fugu Ultra v2 features a context length of 1M tokens. The maximum output length reaches 128K tokens. It accepts both text and image inputs for multi-modal processing. The pricing stands at $5 per 1M input tokens and $30 per 1M output tokens.
| Aspect | Fugu Max | Fugu Ultra v2 |
|---|---|---|
| Input token price per million | $2 | $5 |
| Output token price per million | $6 | $30 |
| Number of top benchmarks | 6 | 5 |
| Context window | Not specified in release | 1M tokens |
| Max output length | Not specified in release | 128K tokens |
| Modalities | Not specified in release | Text and image |
The pricing for Fugu Max stands at $2 per million input tokens and $6 per million output tokens. This represents a 40 to 60 percent reduction compared to Sonnet 5, GPT 5.6 Terra, and Kimi K3. The lower rates apply to the cost-efficient orchestration focus of Fugu Max.
What market and stakeholder implications follow from the Fugu releases?
Lower costs from Fugu Max may encourage wider use of multi-agent systems in enterprise settings. Stakeholders gain options that do not depend on the most expensive closed models. The performance of Fugu Ultra v2 shows that high capability remains achievable through orchestration. Enterprises evaluating AI budgets now have additional data points for cost-benefit analysis.
Platforms such as OpenRouter and Krater.ai provide access to these models. This distribution method increases availability for developers and researchers. The result could shift market dynamics away from single-provider dependence. NVIDIA Nemotron appears among referenced models in related discussions of open pools.
- Fugu Ultra v2 leads on DeepSWE benchmark.
- Fugu Ultra v2 leads on Toolathon benchmark.
- Fugu Ultra v2 leads on GDP.pdf benchmark.
- Fugu Ultra v2 leads on Chartography benchmark.
- Fugu Ultra v2 leads on SWEFish benchmark.
What reactions have experts and the company expressed about these models?
Fugu Max asks: What is the best possible output we can deliver at the lowest possible cost? Fugu Ultra v2 asks: What is the absolute highest capability we can achieve on complex, multi-step tasks?Sakana AI
Sakana AI also states that Fugu Ultra v2 achieves these scores without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool. This underscores the independence from closed frontier models. The statements highlight the intentional design choices in model selection for the agent pool.
What developments might follow for Sakana AI and the field of multi-agent systems?
The dual release indicates Sakana AI will continue refining orchestration methods. Future work may incorporate additional open models or extend support for more modalities. The current results set a new reference point for cost and performance balance. Observers will track whether similar orchestration patterns appear in subsequent releases from other providers.
Industry observers may monitor adoption rates on available APIs. The success on benchmarks like DeepSWE and Toolathon could influence how other providers structure their agent offerings. Overall, the Pareto frontier appears extended through these orchestration techniques. The approach provides evidence that closed frontier models are not required for leading results on selected evaluations.
Frequently asked
How do Fugu Max and Fugu Ultra v2 compare to closed frontier models in performance and cost?
Fugu Ultra v2 reaches top or joint-top scores on five benchmarks without using GPT-6 Astra or Claude Fable 5.1. Fugu Max delivers leading scores on six benchmarks at 40 to 60 percent lower output pricing than listed competitors.
Sources
- Sakana AI — Sakana AI released Fugu Max and Fugu Ultra v2 on September 11, 2026, pushing orchestration forward along both axes simultaneously.
- Sakana AI — Fugu Ultra (v2.0) achieves the best or joint-best score on five of eight benchmarks: GDP.pdf, Chartography, DeepSWE, Toolathon, and SWEFish. Fugu Max prices at $2 per million input tokens and $6 per million output tokens.
- Krater.ai — 1M context, 128K max output — Fugu Ultra v2 context and output limits