分离机器人运动与臂动对视觉的影响,提升移动操作的预测精度。
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

- 通过独立接口解耦相机自身运动与机体动作
- 动作误差降低21.7%,仅用2595万参数实现高效适配
- 适合需要协同控制与移动视角的机器人系统研究
移动操作要求机器人预测足式运动与机械臂动作如何共同影响未来观测和控制。现有世界-动作模型多针对固定基座平台设计,未明确区分相机自运动与基体及臂部动作。本文提出DECOWAM,一种全身式世界-动作模型,通过专用条件接口解耦各因素。该模型冻结经适配的FastWAM主干,训练残差适配器、从特权观测中蒸馏的动作等效未来瓶颈、对抗分离的基体与臂部隐变量,以及基于基体速度的视频预测条件。我们还构建了真实机器人数据集ARMDOG,同步视频、全身状态与动作及语言信息。在固定回放协议下,DECOWAM在视频与动作预测上均优于FastWAM,动作均方误差降低21.7%,仅需2595万可训练参数。在每种方法79次闭环试验中,其展现出最优的全身协调性与基体位移鲁棒性,任务完成率与最强基线相当。结果表明,具身感知的因子分解可在移动视角下实现参数高效的联合视觉预测与全身控制。
原文摘要 · Abstract (English)
Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOWAM, a whole-body world-action model that separates these factors through dedicated conditional interfaces. DECOWAM freezes an adapted FastWAM backbone and trains residual adapters, an action-equivalent future bottleneck distilled from privileged observations, adversarially separated base and arm latents, and base-velocity conditioning for video prediction. We further introduce ARMDOG, a real-robot dataset that synchronizes video, whole-body state and action, and language. On a fixed replay protocol, DECOWAM improved both future-video and action prediction over FastWAM, reducing action MSE by 21.7% with 25.95M trainable adaptation parameters. Across 79 closed-loop trials per method, it achieved the highest observed whole-body coordination and base-displacement robustness among the compared systems, while task completion remained comparable to the strongest baseline. These results show that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。