arXiv:2606.09215cs.RO2026-06被引 2

让机器人用单摄像头实时完成全身动作,比现有方法成功率高30%以上。

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

论文配图:MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation
图 1 · 摘自论文原文
  • 用统一运动潜空间替代上下身分离控制,实现全身协同
  • 在真实机器人上实现实时运行,任务成功率提升超30%
  • 适合需要自然协调动作的复杂人形机器人任务

世界动作模型(WAMs)将视频动态先验与策略结合,在桌面操作任务中表现良好,但对高维视频-动作潜变量进行迭代去噪导致其难以满足人形机器人实时运动操控需求。现有层级范式下,高层策略仅控制上半身,低层控制器跟踪粗略基底指令,使上下半身动作空间不一致,腿部仅用于平衡,限制了整体性能。本文提出MotionWAM,通过单个第一人称摄像头驱动自主人形机器人全身运动,将策略条件于视频世界模型中间去噪特征。它采用统一运动潜空间,联合预测包含行走、躯干运动、高度调节、足部交互和手部操作在内的全身体态动作。通过三阶段学习框架逐步适配视觉动态与目标人形体感。在九项真实单位兔G1任务中,MotionWAM实现实时运行,整体成功率显著优于在同一演示数据上微调的视觉-语言-动作(VLA)基线,提升超过30%,并能执行上/下肢解耦策略无法达成的任务驱动足部交互。结果表明,预训练视频的WAM可从桌面操作扩展至协调、类人式的全身人形控制。

原文摘要 · Abstract (English)

World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on tabletop manipulation, but iterative denoising over high-dimensional video-action latents leaves them too slow for real-time humanoid loco-manipulation. The problem is compounded by the dominant hierarchical paradigm, in which a high-level manipulation policy controls only the upper body while a low-level controller tracks coarse base commands -- placing upper and lower body in inconsistent action spaces and reducing the legs to balance-preserving locomotion. We present MotionWAM, a real-time WAM that drives autonomous humanoid loco-manipulation from a single egocentric camera by conditioning the policy on the intermediate denoising features of a video world model. MotionWAM replaces the upper-lower split with a unified motion latent and predicts whole-body motion tokens that jointly cover locomotion, torso motion, height regulation, foot interaction, and hand manipulation in a single action space. A three-stage learning framework progressively adapts the video world model to egocentric visual dynamics and to the target humanoid embodiment. On nine real-world Unitree G1 tasks, MotionWAM runs in real time, substantially outperforms Vision-Language-Action (VLA) baselines fine-tuned on the same demonstrations by over 30% in overall success rate, and executes task-driven foot interaction that decoupled upper-lower policies cannot reach. Our results suggest that video-pretrained WAMs can be lifted from tabletop manipulation to coordinated, human-like whole-body humanoid control.

人形机器人实时控制动作建模全身协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。