让机器人模型具备全局规划与容错能力,突破传统预测的局限。
HarnessWAM: Bridging Prediction and Deliberation in World Action Models

- 用任务管理器和任务图维持场景认知,实现结构化决策。
- 在RoboMemArena上达成59.6%全任务成功率,子任务成功率69.9%。
- 适合需要长期规划与失败恢复的复杂机器人任务场景。
世界动作模型(WAMs)联合学习环境动态与机器人动作,将物理演化先验引入具身控制。然而,有限时域的预测与动作生成难以应对需全局规划、跨阶段状态保持、执行验证与故障恢复的复杂任务,这种不匹配称为WAM的预测-思辨鸿沟。为此,我们提出HarnessWAM,一种面向WAM的代理式框架。HarnessWAM利用基于视觉语言模型的任务管理器维护基于证据的场景信念与结构化任务图;通过能力条件化的可执行空间投影,将开放语义计划约束为满足任务依赖、具身状态约束及底层WAM能力边界的原子技能序列。执行中,系统采用事件驱动、双时间尺度反馈机制:轻量级进展估计算器提供高频执行证据,任务管理器在关键节点结合当前观测、任务状态与交互历史,判断是否推进、获取新观测、修正计划或启动局部恢复。该机制使机器人可在子任务失败后恢复状态并继续执行,无需丢弃已有场景知识。HarnessWAM在RoboMemArena上实现59.6%的全任务成功率与69.9%的子任务成功率,在RoboCerebra Ideal上达到23.7%的成功率。结果表明,模型外的结构化状态维护与闭环代理决策能有效将WAM的局部控制能力拓展为可规划、可验证、可恢复的具身任务执行。
原文摘要 · Abstract (English)
World Action Models (WAMs) jointly learn environmental dynamics and robot actions, introducing priors over physical evolution into embodied control. However, finite-horizon prediction and action generation are insufficient for complex embodied tasks that require global planning, cross-stage state maintenance, execution verification, and failure recovery. We refer to this mismatch as the prediction-deliberation gap of WAMs. To address this gap, we propose HarnessWAM, an agentic framework for WAMs. HarnessWAM employs a vision-language-model-based Task Manager to maintain an evidence-grounded scene belief and a structured task graph. A capability-conditioned executable-space projection further constrains open-ended semantic plans into sequences of atomic skills that satisfy task dependencies, embodiment-state constraints, and the capability boundary of the underlying WAM. During execution, HarnessWAM operates through an event-driven, dual-timescale feedback loop: a lightweight progress estimator continuously provides high-frequency execution evidence, while the Task Manager deliberates at salient milestones by jointly considering the current observation, task state, and interaction history to determine whether to advance the task, acquire additional observations, revise the plan, or initiate local recovery. This mechanism enables the robot to recover its state after a subtask failure and resume execution without discarding previously acquired scene knowledge. HarnessWAM achieves state-of-the-art full-task and subtask success rates of 59.6% and 69.9% on RoboMemArena, and an SR of 23.7% on RoboCerebra Ideal. These results demonstrate that model-external structured state maintenance and closed-loop agentic decision making can effectively extend the local control capabilities of WAMs into embodied task execution that is plannable, verifiable, and recoverable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。