让仿人机器人在复杂地形上自主导航,通过地形自适应的参考轨迹训练强化学习策略。
Terrain Consistent Reference-Guided RL for Humanoid Navigation Autonomy

- 训练时动态调整参考轨迹,使其贴合真实地形,提升足位规划合理性。
- 在仿真中显著提升参考轨迹跟踪精度,硬件实测实现70米以上闭环自主导航。
- 输出标准SE(2)速度接口,可直接对接主流导航规划器,适合实际部署。
我们提出一种方法,用于训练参考引导的感知强化学习行走策略,使参考轨迹在训练过程中与地形几何保持一致。为兼容标准导航自主架构,我们在强化学习训练循环内合成可控制的SE(2)参考轨迹,将期望步位投影到有效落脚点,并调整摆动腿与质心轨迹以匹配地形。所得到的策略提供一个清晰的SE(2)速度接口,与标准导航规划器兼容。仿真结果表明,环境条件化的参考轨迹显著优于环境无关的参考;在硬件上,我们将该策略与MPC+控制屏障函数规划器集成,在包含崎岖地形和连续台阶的户外环境中,实现了超过70米的长时程闭环自主导航,所有感知与计算均在机载系统完成。
原文摘要 · Abstract (English)
We present a method for training reference-guided, perceptive reinforcement learning locomotion policies for humanoid robots in which reference trajectories are modulated in training to be consistent with terrain geometry. Aiming to deploy our method with standard navigation autonomy infrastructure, we synthesize SE(2)-controllable reference trajectories inside the RL training loop, projecting desired footsteps onto valid footholds and adjusting swing-foot and center-of-mass trajectories to match the terrain. The resulting policy exposes a clean SE(2) velocity interface compatible with standard navigation planners. In simulation, environmentally-conditioned references significantly improve reference tracking performance compared to environment agnostic references. On hardware, we integrate the policy with an MPC + control barrier function planner and demonstrate long-horizon (>70m) closed-loop autonomous navigation on the Unitree G1 through outdoor environments containing rough terrain and consecutive flights of stairs, with all sensing and computation onboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。