Optimus-3让智能体同时具备快速反应和深度思考能力,显著提升复杂任务成功率。
Optimus-3: Dual-Router Aligned Mixture-of-Experts Agent with Dual-Granularity Reasoning-Aware Policy Optimization
- 用知识增强的数据生成法合成高质量推理轨迹,解决数据稀缺问题
- 双路由混合专家架构实现系统1与系统2的动态计算分离,提速4倍以上
- 双粒度奖励机制确保思维过程与结果一致,开放任务成功率达60%
构建能在视觉丰富、动态环境中解决开放性任务的通用智能体仍是具身AI的核心目标。尽管Minecraft已成为重要基准,现有智能体常存在认知能力碎片化问题,缺乏反射式执行(系统1)与反思式推理(系统2)的协同。本文提出Optimus-3,一个在统一框架内有机融合双重能力的通用智能体。为应对三大挑战:首先,针对推理数据稀缺,提出知识增强的数据生成管道,从原始交互轨迹中合成高质量系统2推理轨迹,通过注入领域知识有效抑制幻觉,并发布新数据集OptimusM⁴;其次,为调和双系统计算需求差异,设计双路由对齐的混合专家架构,任务路由实现参数解耦防止干扰,层路由动态调节推理深度,形成系统1的“快路径”与系统2的“深路径”;第三,为激活系统2推理能力,提出双粒度推理感知策略优化算法(DGRPO),通过双粒度密集奖励实现过程-结果联合监督,确保思维与答案一致。大量实验表明,Optimus-3在系统2任务上表现领先(规划+21%、描述+66%、具身问答+76%、定位+3.4×、反思+18%),系统1任务也显著提升(长周期动作+3%),开放任务成功率达60%。
原文摘要 · Abstract (English)
Developing generalist agents capable of solving open-ended tasks in visually rich, dynamic environments remains a core pursuit of embodied AI. While Minecraft has emerged as a compelling benchmark, existing agents often suffer from fragmented cognitive abilities, lacking the synergy between reflexive execution (System 1) and deliberative reasoning (System 2). In this paper, we introduce Optimus-3, a generalist agent that organically integrates these dual capabilities within a unified framework. To achieve this, we address three fundamental challenges. First, to overcome the scarcity of reasoning data, we propose a Knowledge-Enhanced Automated Data Generation Pipeline. It synthesizes high-quality System 2 reasoning traces from raw System 1 interaction trajectories, effectively mitigating hallucinations via injection of domain knowledge. We release the resulting dataset, \textbf{OptimusM$^{4}$}, to the community. Second, to reconcile the dichotomous computational requirements of the dual systems, we design a Dual-Router Aligned MoE Architecture. It employs a Task Router to prevent task interference via parameter decoupling, and a Layer Router to dynamically modulate reasoning depth, creating a computational ``Fast Path'' for System 1 and a ``Deep Path'' for System 2. Third, to activate the reasoning capabilities of System 2, we propose Dual-Granularity Reasoning-Aware Policy Optimization (DGRPO) algorithm. It enforces Process-Outcome Co-Supervision via dual-granularity dense rewards, ensuring consistency between the thought process and the answer. Extensive evaluations demonstrate that Optimus-3 surpasses existing state-of-the-art methods on both System~2 (21$\%$ on Planning, 66\% on Captioning, 76\% on Embodied QA, 3.4$\times$ on Grounding, and 18\% on Reflection) and System~1 (3\% on Long-Horizon Action) tasks, with a notable 60\% success rate on open-ended tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。