arXiv:2409.16784cs.ROcs.LG2024-09ICRA被引 32

用世界模型提升机器人视觉行走的感知能力

World Model-based Perception for Visual Legged Locomotion

  • 构建环境世界模型,用模拟训练驱动真实场景行走
  • 在真实地形上实现比现有方法更优的通行能力和鲁棒性
  • 适合对机器人自主导航与仿真迁移感兴趣的开发者

在复杂地形上实现足式机器人行走极具挑战,需结合本体感知与视觉信息精确理解自身与环境。然而,直接从高维视觉输入学习往往数据效率低且过程复杂。传统方法先利用特权信息训练教师策略,再让学生策略通过视觉模仿其行为,但因输入信息差距导致性能受限,且学习过程不自然。受动物基于对世界的理解自主行走的启发,我们提出一种简单有效的世界模型感知方法(WMP),通过构建环境世界模型并基于该模型训练策略。实验表明,尽管完全在仿真中训练,该世界模型仍能准确预测真实世界轨迹,为控制器提供有效信号。大量仿真与真实世界实验验证,WMP在通行性与鲁棒性上优于当前最优基线方法。视频与代码见:https://wmp-loco.github.io/。

原文摘要 · Abstract (English)

Legged locomotion over various terrains is challenging and requires precise perception of the robot and its surroundings from both proprioception and vision. However, learning directly from high-dimensional visual input is often data-inefficient and intricate. To address this issue, traditional methods attempt to learn a teacher policy with access to privileged information first and then learn a student policy to imitate the teacher's behavior with visual input. Despite some progress, this imitation framework prevents the student policy from achieving optimal performance due to the information gap between inputs. Furthermore, the learning process is unnatural since animals intuitively learn to traverse different terrains based on their understanding of the world without privileged knowledge. Inspired by this natural ability, we propose a simple yet effective method, World Model-based Perception (WMP), which builds a world model of the environment and learns a policy based on the world model. We illustrate that though completely trained in simulation, the world model can make accurate predictions of real-world trajectories, thus providing informative signals for the policy controller. Extensive simulated and real-world experiments demonstrate that WMP outperforms state-of-the-art baselines in traversability and robustness. Videos and Code are available at: https://wmp-loco.github.io/.

足式机器人世界模型仿真迁移视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。