Sakana: Fugu Ultra

by Sakana intelligent reasoning long-context

Learned multi-agent orchestration system from Sakana AI that routes tasks across frontier models (Claude, Gemini, GPT) and can recursively call itself. Leads 6 of 11 benchmarks (vendor-reported) with a 1M-token context window, though performance depends on third-party API availability. Available via Sakana AI API or OpenRouter. Top-tier multi-step reasoning and coding benchmarks 1-million-token context window for long documents Dynamic routing across multiple frontier models Built-in web search and configurable reasoning effort Higher latency vs. base Fugu model Agent pool is fixed, not customizable Performance depends on third-party API uptime Orchestration tokens incur additional real costs Competitive coding and Kaggle challenges Complex multi-step AI research tasks Cybersecurity analysis and penetration testing Advanced agentic workflows with web search

Choose a model to compare against Fugu Ultra

Specifications

Specifications for Fugu Ultra
LabSakana
Context window1,000,000
Input price $5.00/1M
Output price $30.00/1M
Release2026-06-01 00:00:00

Strengths

  • Top-tier multi-step reasoning and coding benchmarks
  • 1-million-token context window for long documents
  • Dynamic routing across multiple frontier models
  • Built-in web search and configurable reasoning effort

Weaknesses

  • Higher latency vs. base Fugu model
  • Agent pool is fixed, not customizable
  • Performance depends on third-party API uptime
  • Orchestration tokens incur additional real costs

Best for

  • Competitive coding and Kaggle challenges
  • Complex multi-step AI research tasks
  • Cybersecurity analysis and penetration testing
  • Advanced agentic workflows with web search

In Depth: Fugu Ultra

Benchmark Performance

Fugu Ultra claims state-of-the-art on 6 of 11 evaluations, achieving 73.7 on SWE-Bench Pro and 93.2 on LiveCodeBench surpassing comparable frontier systems. However, these results are vendor-reported and independent verification has not yet been confirmed.

According to Sakana AI, Fugu Ultra tops Humanity's Last Exam at 50.0, GPQA-D at 95.5, and TerminalBench 2.1 at 82.1. It trails the base Fugu on SciCode and Long-Context Reasoning. These benchmarks are self-reported; users should treat them as preliminary until third-party validation emerges.

Pricing & Value

Input tokens cost $5 per million, output tokens $30 per million with higher rates for contexts above 272K tokens. This places Fugu Ultra at a premium over base Fugu but competitive given its multi-model orchestration.

At $30/M output tokens, Fugu Ultra's price-per-point on LiveCodeBench (93.2) is roughly $0.32 per point competitive with top-tier models like Claude Opus or GPT-5 Turbo on similar benchmarks. However, orchestration overhead can add hidden costs. Base Fugu offers lower latency and no orchestration surcharges.

Who Should Use This

Developers building agentic pipelines, AI researchers validating multi-step reasoning, and cybersecurity analysts needing deep code analysis. Not ideal for latency-sensitive or custom-agent deployments.

  • Competitive coders: top-tier LiveCodeBench scores, but watch latency and orchestration costs
  • AI researchers: leads Humanity's Last Exam, but vendor-reported benchmarks only
  • Cybersecurity analysts: excels in terminal and code tasks, but agent pool is fixed
  • Anti-persona: teams needing low-latency inference or custom model routing should consider base Fugu

Release & Version History

Launched June 22, 2026, Fugu Ultra builds on Sakana AI's multi-agent research published at ICLR 2026. It is a learned orchestration layer, not a traditional monolithic model.

Fugu Ultra follows the earlier base Fugu release, which offered lower latency and optional model opt-outs. Sakana AI released two ICLR 2026 papers and an arXiv technical report (2606.21228) detailing the system's recursive multi-agent architecture. Weights are not distributed; access is API-only via Sakana or OpenRouter. Regional restrictions may apply (unconfirmed EU/EEA limitation reported by third parties).

Sources & Further Reading

Related Models