让机器人通过多智能体协作,突破单一策略限制,实现更鲁棒的决策能力。
EMERGE-Policy: A Robot Mind Emerges Beyond a Single Policy
- 构建图结构多智能体框架,分角色协同完成感知、推理、执行等任务。
- 无需微调即在多个公开基准上表现优异,真实机器人实验验证了可靠性。
- 适合研究复杂机器人系统、自主决策与多智能体协同的学者和工程师。
机器人的有效「心智」不必依赖单一策略,而可在专业化组件共同参与感知、推理、预测、执行、验证与记忆的协同过程中涌现。EMERGE-Policy 将这一理念转化为一种图结构的代理框架,协调能力调用与信息交换。主代理在活动上下文窗口中维护任务级状态,而角色专用子代理在隔离上下文中处理感知、执行监控、验证与记忆整合,并返回结构化、任务相关的证据。角色上下文通过仅暴露决策相关证据来控制信息负荷,功能型技能接口将异构后端组合为操作、想象与评估技能。基于标准的验证、文本故障诊断与分支栈恢复机制实现局部纠错,标记感知的外部记忆保留任务相关状态。它们的闭环交互实现了名为 EMERGE-Policy 的系统级策略。无需额外微调,该系统在多个具有广泛影响的公开基准上取得卓越性能,并完成了系列真实机器人实验。这些结果表明,通过将不同功能子任务分配给多个代理并实现并发协作,以及将模型视为可调用的技能,EMERGE-Policy 能够将鲁棒机器人策略拓展至单次运行之外的持续适应能力。
原文摘要 · Abstract (English)
A robot's effective ``mind'' need not reside in a single policy. It can emerge when specialized components perceive, reason, predict, act, verify, and remember within a shared orchestration process. EMERGE-Policy turns this perspective into a graph-structured agentic framework that coordinates both capability invocation and information exchange. A Main Agent retains task-level state within an active context window, while role-specific Sub Agents process perception, execution monitoring, verification, and memory consolidation in isolated contexts and return structured, task-relevant evidence. Role-specific contexts control information load by exposing only decision-relevant evidence to the Main Agent, while the functional Skill interface composes heterogeneous backends as Operational, Imagination, and Evaluation Skills. Criterion-grounded verification, textual failure diagnosis, and Branch Stack recovery provide localized correction, with token-aware external memory preserving task-relevant state. Together, their closed-loop interaction realizes the system-level policy captured by the name EMERGE-Policy. Without additional fine-tuning, we achieved outstanding performance on several public benchmark that have had a wide-reaching impact, and conducted a series of real robot experiments. These system-level results suggest that through the division of different functional sub-tasks among multiple agents and their concurrent collaboration, as well as the technical paradigm where the model is regarded as a skill and called within the framework, EMERGE-Policy can extend the robust robot policies beyond isolated runs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。