让机器人安全走窄路,分两阶段训练提升稳定性和适应性。
Traversing Narrow Paths: A Two-Stage Reinforcement Learning Framework for Robust and Safe Humanoid Walking
- 先用物理模板规划脚位,再用强化学习微调修正,分步优化
- 在0.2米宽、3米长的梁上20次全成功,零失败
- 适合需要高精度足部控制的仿人机器人场景
在狭窄路径上行走对仿人机器人而言极具挑战,因其所需落脚点稀疏且安全要求高。纯基于模板或端到端强化学习的方法在此类地形下表现不佳。本文提出一种两阶段训练框架:第一阶段结合物理模板足位规划器与低层足位跟踪器;第二阶段引入轻量级感知辅助的足位修正模块。通过从平坦地面逐步过渡到窄路的课程训练,最终控制器能稳健追踪并安全修正目标落脚点,实现窄路精准落脚。该框架兼顾物理模型的可解释性与强化学习的泛化能力,支持高效仿真到现实迁移。实验表明,所学策略在成功率、中心线跟随和安全裕度方面均优于纯模板或纯强化学习基线。在Unitree G1机器人上验证,连续20次成功穿越宽0.2米、长3米的横梁,无一失败。
原文摘要 · Abstract (English)
Traversing narrow paths is challenging for humanoid robots due to the sparse and safety-critical footholds required. Purely template-based or end-to-end reinforcement learning-based methods suffer from such harsh terrains. This paper proposes a two stage training framework for such narrow path traversing tasks, coupling a template-based foothold planner with a low-level foothold tracker from Stage-I training and a lightweight perception aided foothold modifier from Stage-II training. With the curriculum setup from flat ground to narrow paths across stages, the resulted controller in turn learns to robustly track and safely modify foothold targets to ensure precise foot placement over narrow paths. This framework preserves the interpretability from the physics-based template and takes advantage of the generalization capability from reinforcement learning, resulting in easy sim-to-real transfer. The learned policies outperform purely template-based or reinforcement learning-based baselines in terms of success rate, centerline adherence and safety margins. Validation on a Unitree G1 humanoid robot yields successful traversal of a 0.2m wide and 3m long beam for 20 trials without any failure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。