arXiv:2606.21228cs.LG2026-06被引 2

让多个大模型协作解题,性能超越单个模型。

Sakana Fugu Technical Report

论文配图:Sakana Fugu Technical Report
图 1 · 摘自论文原文
  • 用语言模型动态构建任务支架,协调多个AI agent协同工作。
  • 在6项挑战性任务中达到当前公开模型最优表现,如SWE-Bench Pro和GPQA-Diamond。
  • 适合研究多智能体系统与动态协作架构的开发者和研究人员。

前沿大语言模型能力持续提升,不同厂商逐渐聚焦特定领域。这引出一个新目标:如何将多个大模型的专业能力整合为集体智能系统。为此,我们报告了Sakana Fugu系列编排模型的开发,该模型能调用并放大一组LLM代理的能力。Fugu模型自身是语言模型,可理解用户问题,并动态构建代理结构来解决任务。通过这些自适应结构,Fugu实现了超越任何单一模型的性能,在多项挑战性任务上达到当前公开模型最优,包括SWE-Bench Pro、Terminal Bench、LiveCodeBench、GPQA-Diamond、Humanity's Last Exam和CharXiv Reasoning。我们发布了两个版本:Fugu,兼顾性能与延迟,适合日常使用;Fugu-Ultra,优先保障最难问题的答案质量。本文描述了涵盖大规模微调、进化算法和强化学习的训练范式,以及支撑其落地生产的基础设施与核心设计原则。希望本报告推动对多智能体系统与动态查询适配代理结构的研究,探索通过集体智能迈向下一代AI能力的新路径。

原文摘要 · Abstract (English)

The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains. This raises a natural next objective: how to combine the individual specializations of various LLMs into a collectively intelligent system. To this end, we report the development of Sakana Fugu, a family of orchestrator models that harness and amplify the capabilities of an LLM agent team. Fugu models are themselves language models trained to understand user queries and dynamically devise agentic scaffolds to solve them. Through these adaptive scaffolds, Fugu accesses performance beyond any individual LLM agent, achieving state-of-the-art results compared to other publicly accessible models across a range of challenging tasks, including SWE-Bench Pro, Terminal Bench, LiveCodeBench, GPQA-Diamond, Humanity's Last Exam, and CharXiv Reasoning. We release two models: Fugu, which balances performance with latency for everyday use, and Fugu-Ultra, which prioritizes answer quality on the hardest problems. We describe our training paradigm, which encompasses large-scale fine-tuning, evolutionary algorithms, and reinforcement learning approaches, along with the infrastructure and core design principles that turn these methods into a production system. We hope this report encourages further research into multi-agent systems and dynamic, query-adaptive agentic scaffolds as a path toward the next frontier of AI capabilities, accessed through collective intelligence.

多智能体编排模型集体智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。