Sakana: Fugu Ultra
Learned multi-agent orchestration system from Sakana AI that routes tasks across frontier models (Claude, Gemini, GPT) and can recursively call itself. Leads 6 of 11 benchmarks (vendor-reported) with a 1M-token context window, though performance depends on third-party API availability. Available via Sakana AI API or OpenRouter.
Specifications
| Attribute | Value |
|---|---|
| Lab | Sakana |
| Tags | Intelligent Reasoning Long Context |
| Release Date | 2026-06 |
| Context Window | 1,000,000 tokens |
| Input Price / 1M | $5.00 |
| Output Price / 1M | $30.00 |
| Input Modalities | Text, Image |
| Output Modalities | Text |
Strengths
- Top-tier multi-step reasoning and coding benchmarks
- 1-million-token context window for long documents
- Dynamic routing across multiple frontier models
- Built-in web search and configurable reasoning effort
Weaknesses
- Higher latency vs. base Fugu model
- Agent pool is fixed, not customizable
- Performance depends on third-party API uptime
- Orchestration tokens incur additional real costs
Best For
- Competitive coding and Kaggle challenges
- Complex multi-step AI research tasks
- Cybersecurity analysis and penetration testing
- Advanced agentic workflows with web search
In Depth: Fugu Ultra
Benchmark Performance
Fugu Ultra claims state-of-the-art on 6 of 11 evaluations, achieving 73.7 on SWE-Bench Pro and 93.2 on LiveCodeBench surpassing comparable frontier systems. However, these results are vendor-reported and independent verification has not yet been confirmed.
According to Sakana AI, Fugu Ultra tops Humanity's Last Exam at 50.0, GPQA-D at 95.5, and TerminalBench 2.1 at 82.1. It trails the base Fugu on SciCode and Long-Context Reasoning. These benchmarks are self-reported; users should treat them as preliminary until third-party validation emerges.
Pricing & Value
Input tokens cost $5 per million, output tokens $30 per million with higher rates for contexts above 272K tokens. This places Fugu Ultra at a premium over base Fugu but competitive given its multi-model orchestration.
At $30/M output tokens, Fugu Ultra's price-per-point on LiveCodeBench (93.2) is roughly $0.32 per point competitive with top-tier models like Claude Opus or GPT-5 Turbo on similar benchmarks. However, orchestration overhead can add hidden costs. Base Fugu offers lower latency and no orchestration surcharges.
Who Should Use This
Developers building agentic pipelines, AI researchers validating multi-step reasoning, and cybersecurity analysts needing deep code analysis. Not ideal for latency-sensitive or custom-agent deployments.
- Competitive coders: top-tier LiveCodeBench scores, but watch latency and orchestration costs
- AI researchers: leads Humanity's Last Exam, but vendor-reported benchmarks only
- Cybersecurity analysts: excels in terminal and code tasks, but agent pool is fixed
- Anti-persona: teams needing low-latency inference or custom model routing should consider base Fugu
Release & Version History
Launched June 22, 2026, Fugu Ultra builds on Sakana AI's multi-agent research published at ICLR 2026. It is a learned orchestration layer, not a traditional monolithic model.
Fugu Ultra follows the earlier base Fugu release, which offered lower latency and optional model opt-outs. Sakana AI released two ICLR 2026 papers and an arXiv technical report (2606.21228) detailing the system's recursive multi-agent architecture. Weights are not distributed; access is API-only via Sakana or OpenRouter. Regional restrictions may apply (unconfirmed EU/EEA limitation reported by third parties).