arXiv:2608.00725cs.RO2026-08被引 1

让机器人预测动作后果更精准,兼顾速度与真实感。

SelfWAM: A Self-Grounded Unified World Action Model for Fast Robot Control

论文配图:SelfWAM: A Self-Grounded Unified World Action Model for Fast Robot Control
图 1 · 摘自论文原文
  • 用动作自身信息直接引导未来视觉预测,避免泛化偏差。
  • 在真实任务中提升策略性能,同时保持快速推理速度。
  • 适合需要高精度动作反馈的机器人控制场景。

世界动作模型(WAM)通过联合建模动作与未来观测来提升机器人策略学习效果。然而,仅以任务提示和观测上下文作为未来预测条件,可能捕捉到通用的任务进展而非具体动作带来的后果。本文提出SelfWAM,一种基于模态专用混合变压器(MoT)架构的统一自锚定世界动作模型,能联合预测动作、动作相关的未来RGB图像及机器人自掩码,从而将未来预测锚定在机器人可见身体及其动作引发的运动上。在联合训练中,未来视觉查询可访问演示动作的干净副本,使视频分支成为动作特异性后果模型,同时保持快速的动作推理路径不变。为聚焦视频学习于动作相关视觉变化,采用任务特定目标进行未来机器人自掩码预测,去除外观细节,其时间演化与动作紧密耦合。清洁动作条件与未来自掩码监督共同使未来预测更直接反映执行动作对机器人可见运动及周围场景的影响。在RoboTwin 2.0和真实世界操作任务上的实验表明,SelfWAM生成更具动作敏感性的未来预测,维持快速策略推断,同时提升策略性能。

原文摘要 · Abstract (English)

World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future observations. However, conditioning future prediction only on the task prompt and observation context risks capturing generic task progression rather than the action-specific consequences of the executed action. We introduce SelfWAM, a unified self-grounded WAM built on a modality-specialized Mixture-of-Transformers (MoT) architecture that jointly predicts actions, action-conditioned future RGB frames, and robot self-masks, thereby grounding future prediction in the robot's visible body and its action-induced motion. During joint training, SelfWAM allows future visual queries to attend to a clean copy of the demonstrated action, turning the video branch into an action-specific consequence model while leaving the fast action-only inference path unchanged. To focus video learning on action-relevant visual changes, we use prompt-specific objectives for future robot self-mask prediction, which removes appearance details and provides a target whose temporal evolution is tightly coupled with the conditioning action. Together, clean-action conditioning and future self-mask supervision make future predictions more directly reflect how the executed action changes the robot's visible motion and the surrounding scene. Experiments on RoboTwin 2.0 and real-world manipulation tasks show that SelfWAM produces more action-sensitive futures and preserves fast policy inference, while improving policy performance.

机器人控制动作预测自锚定视觉建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。