arXiv:2606.30616cs.CL2026-06被引 7

350亿参数的智能体模型,靠扩展任务视野而非增大参数量,达到万亿级模型性能。

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

论文配图:Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
图 1 · 摘自论文原文
  • 通过扩展智能体的任务规划长度和能力多样性,实现长周期任务处理。
  • 在多个长程任务基准上超越1万亿参数模型,平均轨迹达4.5万词元。
  • 适合追求高效长任务智能体的开发者与研究者参考使用。

我们提出Agents-A1,一个350亿参数的专家混合型智能体模型,通过扩展智能体视野而非增加参数量,实现了万亿级参数模型的性能。从两个角度探索智能体视野的扩展:长周期任务轨迹与异构能力的提升。为此构建了连接外部知识、动作、观测与验证结果的长周期知识-行动基础设施,生成平均长度为45,000词元的智能体轨迹。基于此,采用三阶段训练方案:首先进行全领域监督微调以对齐基础模型的广泛智能体行为;其次训练各领域的专用教师模型以捕获专业能力;最后提出多教师域路由在线蒸馏与显著词汇对齐方法,提升跨域知识迁移效率,将六个异构领域统一为可部署的学生模型。Agents-A1在多个长周期任务基准测试中表现强劲,优于如Kimi-K2.6和DeepSeek-V4-pro等1万亿参数模型,在SEAL-0(56.4)、IFBench(80.6)、HiPhO(46.4)、FrontierScience-Olympiad(79.0)和MolBench-Bind(56.8)上取得领先成绩,并在SciCode(44.3)、HLE(47.6)和BrowseComp(75.5)上保持高度竞争力。本工作为使用350亿参数智能体实现或逼近万亿级模型在长周期任务中的表现提供了实用路径。

原文摘要 · Abstract (English)

We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories and scaling heterogeneous agent abilities. To support this goal, we build a long-horizon knowledge-action infrastructure that connects external knowledge, actions, observations, and verifier outcomes, producing agentic trajectories with an average length of 45K tokens. Based on this, we train Agents-A1 with a three-stage recipe. First, we perform full-domain supervised fine-tuning to align the base model with broad agentic behaviors. Second, we train domain-level teacher models to capture specialized expertise in each domain. Third, we propose a multi-teacher domain-routed on-policy distillation with salient vocabulary alignment to improve knowledge transfer efficiency across different domains, unifying six heterogeneous domains into one deployable student model. Agents-A1 achieves strong and broad performance for long-horizon agent benchmarks. Compared with 1T-parameter model such as Kimi-K2.6 and DeepSeek-V4-pro, Agents-A1 achieves leading results on SEAL-0 (56.4), IFBench (80.6), HiPhO (46.4), FrontierScience-Olympiad (79.0), and MolBench-Bind (56.8), and remains highly competitive on SciCode (44.3), HLE (47.6) and BrowseComp (75.5). We hope this work provides the community with a practical path for scaling the horizon using a 35B agent that can reach or match the performance of 1T models on long-horizon tasks.

智能体长周期任务模型压缩专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。