{"id":324,"date":"2026-06-28T08:12:50","date_gmt":"2026-06-28T00:12:50","guid":{"rendered":"https:\/\/aidashxp.com\/sakana-fugu-ultra-review\/"},"modified":"2026-06-28T08:12:50","modified_gmt":"2026-06-28T00:12:50","slug":"sakana-fugu-ultra-review","status":"publish","type":"post","link":"https:\/\/aidashxp.com\/en\/sakana-fugu-ultra-review\/","title":{"rendered":"Sakana Fugu Ultra in-depth review: Japanese AI orchestration system, SWE-Bench Pro scores 73.7 to challenge cutting-edge models"},"content":{"rendered":"<p class=\"wp-block-paragraph\">Japan AI Lab<strong>Sakana AI<\/strong>officially released its latest product on June 22<strong>Fugu and Fugu Ultra<\/strong>\u2014\u2014This is not a traditional large language model, but a<strong>Multi-agent orchestration system<\/strong>. It dynamically schedules an \"agent pool\" composed of multiple cutting-edge models through a single API interface to automatically select the best execution strategy for complex tasks. Fugu Ultra achieved a score of 73.7 on the SWE-Bench Pro programming benchmark, and Sakana AI claims its performance is comparable to Anthropic's Mythos and Fable 5-level models.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">core competencies<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Fugu\u2019s core innovation lies in its<strong>\u201cOrchestration as model\u201d<\/strong>architecture. Users only need to send a request to a single endpoint, and Fugu will automatically determine whether to answer simple questions directly or form a \"team of experts\" to handle complex tasks. The entire process is completely transparent to developers.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Multi-model dynamic scheduling<\/strong>: Fugu itself is a well-trained language model that specifically learns how to call other LLMs, including recursively calling its own instances.<\/li>\n<li><strong>Dual version design<\/strong>: Fugu (Basic Edition) and Fugu Ultra (Ultimate Edition), accessed through the same OpenAI compatible API<\/li>\n<li><strong>Replaceable agent pool<\/strong>: The underlying \"expert model pool\" is completely replaceable, and Sakana plans to add open source models and self-developed models in the future.<\/li>\n<li><strong>Outstanding performance on benchmarks<\/strong>: Fugu Ultra achieved 73.7 points in SWE-Bench Pro and 82.1 points in TerminalBench 2.1, exceeding GPT-5.5 and<a href=\"https:\/\/claude.ai\" target=\"_blank\" rel=\"nofollow noopener\">Claude<\/a> Opus 4.8<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">User experience and limitations<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Judging from early user feedback, Fugu Ultra\u2019s capabilities are indeed impressive, especially when<strong>Programming and Scientific Reasoning<\/strong>Excellent performance in tasks. Sakana AI demonstrates on its product page that Fugu Ultra continues to outperform GPT-5.5, Gemini 3.1 Pro and<a href=\"https:\/\/claude.ai\" target=\"_blank\" rel=\"nofollow noopener\">Claude<\/a> Experimental results on Opus 4.8.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But Fugu also has obvious shortcomings:<strong>API response is slow<\/strong>\u2014\u2014Because it needs to schedule multiple models to work together internally, simple queries also have a fixed orchestration overhead of about 1260 tokens, and the total token consumption of complex queries can reach 5 to 12 times the visible output. In terms of pricing, Fugu Ultra monthly<strong>$200<\/strong>, some users joked, \"It cost 200 US dollars, but it was actually used for less than 3 hours a week.\" also,<strong>EU and EEA are not currently supported<\/strong>visit.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Overall Score<\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table>\n<thead><tr><th>\u7ef4\u5ea6<\/th><th>Score<\/th><th>evaluate<\/th><\/tr><\/thead>\n<tbody>\n<tr><td>functional completeness<\/td><td>8.5 \/ 10<\/td><td>The multi-agent orchestration architecture is unique, but the currently available expert model pool is limited in scope.<\/td><\/tr>\n<tr><td>\u6613\u7528\u6027<\/td><td>8.0 \/ 10<\/td><td>A single API interface reduces the difficulty of integration, but slow response speed affects the experience<\/td><\/tr>\n<tr><td>Cost-effectiveness<\/td><td>7.0 \/ 10<\/td><td>The pricing of US$200\/month is on the high side, and the orchestration overhead causes the actual token cost to be 5-12 times the visible output.<\/td><\/tr>\n<tr><td>\u4e2d\u6587\u652f\u6301<\/td><td>7.5 \/ 10<\/td><td>The underlying model pool contains multi-language models, but is not natively optimized for Chinese<\/td><\/tr>\n<tr><td>\u8f93\u51fa\u8d28\u91cf<\/td><td>8.8 \/ 10<\/td><td>Excellent performance in programming and reasoning tasks, with SWE-Bench Pro score of 73.7 verifying its capabilities<\/td><\/tr>\n<\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Overall rating: 8.0\/10<\/strong><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Summarize<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Sakana Fugu represents a new AI product paradigm - it is no longer about \"training a stronger model\", but<strong>\"Training a Smarter Dispatcher\"<\/strong>. This idea is technically forward-looking and is especially suitable for complex scenarios that require flexible use of the advantages of different models. However, Fugu is currently limited by response speed and pricing, and is more suitable for professional developers who are not cost-sensitive and pursue the ultimate output quality. With the expansion and optimization of the underlying agent pool, Fugu's \"orchestration as model\" concept may become an important direction for the next generation of AI infrastructure.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/aidashxp.com\/en\/\">AI Dash<\/a> \u2014 Discover the best AI tools and get the latest AI model reviews.<\/p>","protected":false},"excerpt":{"rendered":"<p>\u65e5\u672cAI\u5b9e\u9a8c\u5ba4Sakana AI\u4e8e6\u67082 [&hellip;]<\/p>\n","protected":false},"author":0,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[6],"tags":[],"class_list":["post-324","post","type-post","status-publish","format-standard","hentry","category-ai-coding"],"_links":{"self":[{"href":"https:\/\/aidashxp.com\/en\/wp-json\/wp\/v2\/posts\/324","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/aidashxp.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/aidashxp.com\/en\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/aidashxp.com\/en\/wp-json\/wp\/v2\/comments?post=324"}],"version-history":[{"count":0,"href":"https:\/\/aidashxp.com\/en\/wp-json\/wp\/v2\/posts\/324\/revisions"}],"wp:attachment":[{"href":"https:\/\/aidashxp.com\/en\/wp-json\/wp\/v2\/media?parent=324"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/aidashxp.com\/en\/wp-json\/wp\/v2\/categories?post=324"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/aidashxp.com\/en\/wp-json\/wp\/v2\/tags?post=324"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}