Japan AI LabSakana AIofficially released its latest product on June 22Fugu and Fugu Ultra——This is not a traditional large language model, but aMulti-agent orchestration system. It dynamically schedules an "agent pool" composed of multiple cutting-edge models through a single API interface to automatically select the best execution strategy for complex tasks. Fugu Ultra achieved a score of 73.7 on the SWE-Bench Pro programming benchmark, and Sakana AI claims its performance is comparable to Anthropic's Mythos and Fable 5-level models.
core competencies
Fugu’s core innovation lies in its“Orchestration as model”architecture. Users only need to send a request to a single endpoint, and Fugu will automatically determine whether to answer simple questions directly or form a "team of experts" to handle complex tasks. The entire process is completely transparent to developers.
- Multi-model dynamic scheduling: Fugu itself is a well-trained language model that specifically learns how to call other LLMs, including recursively calling its own instances.
- Dual version design: Fugu (Basic Edition) and Fugu Ultra (Ultimate Edition), accessed through the same OpenAI compatible API
- Replaceable agent pool: The underlying "expert model pool" is completely replaceable, and Sakana plans to add open source models and self-developed models in the future.
- Outstanding performance on benchmarks: Fugu Ultra achieved 73.7 points in SWE-Bench Pro and 82.1 points in TerminalBench 2.1, exceeding GPT-5.5 andClaude Opus 4.8
User experience and limitations
Judging from early user feedback, Fugu Ultra’s capabilities are indeed impressive, especially whenProgramming and Scientific ReasoningExcellent performance in tasks. Sakana AI demonstrates on its product page that Fugu Ultra continues to outperform GPT-5.5, Gemini 3.1 Pro andClaude Experimental results on Opus 4.8.
But Fugu also has obvious shortcomings:API response is slow——Because it needs to schedule multiple models to work together internally, simple queries also have a fixed orchestration overhead of about 1260 tokens, and the total token consumption of complex queries can reach 5 to 12 times the visible output. In terms of pricing, Fugu Ultra monthly$200, some users joked, "It cost 200 US dollars, but it was actually used for less than 3 hours a week." also,EU and EEA are not currently supportedvisit.
Overall Score
| 维度 | Score | evaluate |
|---|---|---|
| functional completeness | 8.5 / 10 | The multi-agent orchestration architecture is unique, but the currently available expert model pool is limited in scope. |
| 易用性 | 8.0 / 10 | A single API interface reduces the difficulty of integration, but slow response speed affects the experience |
| Cost-effectiveness | 7.0 / 10 | The pricing of US$200/month is on the high side, and the orchestration overhead causes the actual token cost to be 5-12 times the visible output. |
| 中文支持 | 7.5 / 10 | The underlying model pool contains multi-language models, but is not natively optimized for Chinese |
| 输出质量 | 8.8 / 10 | Excellent performance in programming and reasoning tasks, with SWE-Bench Pro score of 73.7 verifying its capabilities |
Overall rating: 8.0/10
Summarize
Sakana Fugu represents a new AI product paradigm - it is no longer about "training a stronger model", but"Training a Smarter Dispatcher". This idea is technically forward-looking and is especially suitable for complex scenarios that require flexible use of the advantages of different models. However, Fugu is currently limited by response speed and pricing, and is more suitable for professional developers who are not cost-sensitive and pursue the ultimate output quality. With the expansion and optimization of the underlying agent pool, Fugu's "orchestration as model" concept may become an important direction for the next generation of AI infrastructure.
AI Dash — Discover the best AI tools and get the latest AI model reviews.
