让机器人在稀疏落脚点上敏捷行走,靠强化学习实现精准踩踏。
BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds

- 用采样奖励与双评判器平衡密集与稀疏奖励学习
- 两阶段训练:先平地预训练再真实地形微调
- 结合激光地图实现实时部署,抗干扰能力强
在稀疏落脚点上穿越危险地形对人形机器人构成重大挑战,需精确落脚与稳定行走。现有基于学习的方法因落脚奖励稀疏和学习效率低而表现不佳。为此,我们提出BeamDojo,一种用于稀疏落脚点上敏捷人形行走的强化学习框架。该框架引入针对多边形足部的采样式落脚奖励,并采用双评判器机制,平衡密集行走奖励与稀疏落脚奖励之间的学习过程。为促进充分试错探索,采用两阶段强化学习:第一阶段在平坦地形上训练,同时提供任务-地形感知观测;第二阶段在实际任务地形上微调策略。此外,我们实现了基于机载激光雷达的高程图,支持真实世界部署。大量仿真与真实世界实验表明,BeamDojo在仿真中实现高效学习,在真实环境中实现精准落脚的敏捷行走,即使在显著外部扰动下仍保持高成功率。
原文摘要 · Abstract (English)
Traversing risky terrains with sparse footholds poses a significant challenge for humanoid robots, requiring precise foot placements and stable locomotion. Existing learning-based approaches often struggle on such complex terrains due to sparse foothold rewards and inefficient learning processes. To address these challenges, we introduce BeamDojo, a reinforcement learning (RL) framework designed for enabling agile humanoid locomotion on sparse footholds. BeamDojo begins by introducing a sampling-based foothold reward tailored for polygonal feet, along with a double critic to balancing the learning process between dense locomotion rewards and sparse foothold rewards. To encourage sufficient trial-and-error exploration, BeamDojo incorporates a two-stage RL approach: the first stage relaxes the terrain dynamics by training the humanoid on flat terrain while providing it with task-terrain perceptive observations, and the second stage fine-tunes the policy on the actual task terrain. Moreover, we implement a onboard LiDAR-based elevation map to enable real-world deployment. Extensive simulation and real-world experiments demonstrate that BeamDojo achieves efficient learning in simulation and enables agile locomotion with precise foot placement on sparse footholds in the real world, maintaining a high success rate even under significant external disturbances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。