用预测未来视觉的模型提升人形机器人在复杂地形上的行走稳定性。
World-Model-Augmented Visual Locomotion for Humanoids on Foothold-Constrained Terrain
- 通过联合训练世界模型与策略,利用单帧深度图预测未来状态。
- 在跨越间隙和石块上成功率100%,比基线大幅领先。
- 适合研究机器人自主导航与视觉预测的工程师和学者。
在足点受限地形(如跨步石、缝隙、窄台阶)中,一步失误往往难以挽回,仅依赖当前可见地形的策略易失败。本文提出世界模型增强的视觉行走方法(WM-LOCO),联合训练递归世界模型与PPO策略。基于本体感知和单张机载深度图像,世界模型生成预测性递归特征,指导策略决策,无需显式足点标注。仿真中,WM-LOCO在跨越缝隙和石块任务上成功率达100%,而匹配基线完全失败;在台阶任务上达到基线成功率,同时提升步态效率并降低骨盆加速度。将同一策略部署于物理单位树G1人形机器人,仅使用本体感知和单路深度流,在三类地形上平均成功率达93.3%。
原文摘要 · Abstract (English)
Foothold-constrained terrain is characterized by sparse, discontinuous, or geometrically restricted feasible foot contacts, as encountered on stepping stones, across gaps, and on narrow stair treads. On such terrain, a single misstep often leaves little room to recover, so policies that base foot-placement decisions primarily on the immediately visible terrain are prone to failure. We ask whether a learned predictive summary of near-future observations and rewards can provide the anticipatory information required in such settings. We present World-Model-Augmented Visual Locomotion (WM-LOCO), which jointly trains a recurrent world model and a PPO policy. Conditioned on proprioception and a single onboard depth image, the world model produces a predictive recurrent feature that guides the policy, without explicit foothold labels. In simulation, WM-LOCO succeeds on gaps and stepping stones where a matched baseline fails completely, and matches the baseline's success rate on stairs while improving stride efficiency and reducing pelvis acceleration. We deploy the same policy onboard a physical Unitree G1 humanoid using onboard proprioception and a single depth stream; it traverses all three terrain classes with an average success rate of 93.3%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。