arXiv:2602.03002cs.RO2026-02被引 15

让机器人在复杂地形上稳健行走并负重,支持多方向移动。

RPL: Learning Robust Humanoid Perceptive Locomotion on Challenging Terrains

  • 分两阶段训练:先学不同地形技能,再融合为统一策略。
  • 实测可在20°斜坡、30厘米台阶等挑战地形上负重2公斤行走。
  • 创新深度处理技术,提升对不对称和未知宽度地形的适应力。

人形机器人感知行走取得显著进展,但在复杂地形上实现稳健的多方向行走仍待探索。为此,我们提出RPL,一种两阶段训练框架,实现复杂地形上的多方向行走且具备负载鲁棒性。RPL首先利用高程图观测训练各地形专用专家策略,掌握解耦的行走与操作技能;随后将其蒸馏为基于多视角深度相机的Transformer策略。蒸馏过程中引入两项关键技术:基于速度指令的深度特征缩放,以及随机侧向遮蔽,以应对非对称深度观测和未见地形宽度。为实现高效深度蒸馏,开发了高效的多深度系统,在大规模并行环境中同时对动态机器人网格和静态地形网格进行射线投射,相比现有模拟器的深度渲染管道提速5倍,同时建模真实传感器延迟、噪声与丢包。大量真实世界实验表明,该方法可在20°斜坡、步长分别为22厘米、25厘米、30厘米的楼梯,以及60厘米间距的25厘米×25厘米踏石等挑战地形上,实现带2公斤负载的稳健多方向行走。

原文摘要 · Abstract (English)

Humanoid perceptive locomotion has made significant progress and shows great promise, yet achieving robust multi-directional locomotion on complex terrains remains underexplored. To tackle this challenge, we propose RPL, a two-stage training framework that enables multi-directional locomotion on challenging terrains, and remains robust with payloads. RPL first trains terrain-specific expert policies with privileged height map observations to master decoupled locomotion and manipulation skills across different terrains, and then distills them into a transformer policy that leverages multiple depth cameras to cover a wide range of views. During distillation, we introduce two techniques to robustify multi-directional locomotion, depth feature scaling based on velocity commands and random side masking, which are critical for asymmetric depth observations and unseen widths of terrains. For scalable depth distillation, we develop an efficient multi-depth system that ray-casts against both dynamic robot meshes and static terrain meshes in massively parallel environments, achieving a 5-times speedup over the depth rendering pipelines in existing simulators while modeling realistic sensor latency, noise, and dropout. Extensive real-world experiments demonstrate robust multi-directional locomotion with payloads (2kg) across challenging terrains, including 20° slopes, staircases with different step lengths (22 cm, 25 cm, 30 cm), and 25 cm by 25 cm stepping stones separated by 60 cm gaps.

人形机器人路径规划强化学习仿真加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。