arXiv:2502.16230cs.ROcs.LG2025-02被引 14

让机器人在复杂地形上自主行走,靠重建环境状态实现无感导航。

Learning Humanoid Locomotion with World Model Reconstruction

  • 用端到端模型重建环境状态,驱动步态策略
  • 实测完成3.2公里野外徒步,穿越冰雪与滑地
  • 适合研究真实场景下机器人自主运动的团队

人类形态机器人旨在通过双腿在人类可访问环境中行进。然而,传统研究多集中于受控实验室环境,导致在复杂现实地形中构建控制策略的能力不足。这主要源于传感器数据的局限性和噪声,阻碍了机器人对自身及环境的理解。本文提出世界模型重建(World Model Reconstruction, WMR),一种面向盲视人类形态机器人在挑战性地形中行走的端到端学习方法。我们训练一个估计器显式重构世界状态,并利用该信息增强步态策略。步态策略仅接收重构后的信息作为输入。策略与估计器联合训练,但其间梯度被有意切断,确保估计器专注于世界重建,不受步态策略更新影响。我们在真实世界中评估模型,涵盖粗糙、可变形和滑腻表面,展示了出色的适应性与抗干扰能力。机器人成功完成3.2公里无人干预徒步,掌握了覆盖冰与雪的地形。

原文摘要 · Abstract (English)

Humanoid robots are designed to navigate environments accessible to humans using their legs. However, classical research has primarily focused on controlled laboratory settings, resulting in a gap in developing controllers for navigating complex real-world terrains. This challenge mainly arises from the limitations and noise in sensor data, which hinder the robot's understanding of itself and the environment. In this study, we introduce World Model Reconstruction (WMR), an end-to-end learning-based approach for blind humanoid locomotion across challenging terrains. We propose training an estimator to explicitly reconstruct the world state and utilize it to enhance the locomotion policy. The locomotion policy takes inputs entirely from the reconstructed information. The policy and the estimator are trained jointly; however, the gradient between them is intentionally cut off. This ensures that the estimator focuses solely on world reconstruction, independent of the locomotion policy's updates. We evaluated our model on rough, deformable, and slippery surfaces in real-world scenarios, demonstrating robust adaptability and resistance to interference. The robot successfully completed a 3.2 km hike without any human assistance, mastering terrains covered with ice and snow.

人形机器人自主导航世界建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。