让机器人边走边干活,一模型搞定全身协调动作预测。
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

- 用潜空间预测未来视觉特征,结合扩散模型生成全身动作
- 在11个真实家务任务中表现超越现有方法,动作更平滑自然
- 适合需要边移动边操作的具身智能研究者使用
人类家庭任务常需同时完成行走与操作,要求机器人协调移动、调整姿态、保持平衡并操控物体。现有方法多将运动与操作分开处理,近期世界-动作模型则或聚焦手臂、或依赖视频重建。我们提出 $ω$-0,一种用于真实世界人形机器人协同运动-操作的潜空间预测全身体动作模型。在语言指令、当前视觉观测和机器人本体感知状态下,直接预测可被控制器执行的全身动作潜变量。不同于重建未来视频,$ω$-0以紧凑的未来观测嵌入作为轻量级预测目标,将潜视觉前瞻与基于扩散模型的全身体动作生成相耦合。模型支持第一视角RGB、第二视角RGB及深度输入,并利用基于控制器的仿真回放,将人类/公开数据中的视觉-运动先验转化为机器人可执行的动作潜变量。我们进一步构建了 $ω$-HOME 数据集,包含40+小时真实家庭场景下同步的多视角观测、全身SMPL运动、机器人状态及动作潜变量。在11个真实家庭任务上的实验表明,单一 $ω$-0 模型能生成流畅的边走边操作行为,且持续优于代表性模仿学习、视觉语言动作(VLA)、人形机器人及世界-动作模型(WAM)基线。
原文摘要 · Abstract (English)
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present $ω$-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Given a language instruction, current visual observation, and robot proprioceptive state, $ω$-0 directly predicts controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future videos, $ω$-0 learns compact future observation embeddings as a lightweight predictive objective, coupling latent visual foresight with diffusion-based whole-body action generation. The model supports egocentric RGB, exocentric RGB, and exocentric depth inputs, and leverages controller-based simulation replay to ground human/public visual-motion priors into robot-executable action latents. We further collect $ω$-HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents. Real-world experiments on 11 household tasks demonstrate that a single $ω$-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。