用选择性对抗运动先验,让机器人学会五种步态
Multi-Gait Learning for Humanoid Robots Using Reinforcement Learning with Selective Adversarial Motion Prior

- 对稳定步态用对抗先验加速学习,动态步态则不用避免过度约束
- 相比统一使用对抗先验,收敛更快、误差更低、成功率更高
- 适合想统一训练多种步态的机器人研究者
在统一强化学习框架中为类人机器人学习多样化行走技能仍具挑战,因不同步态对稳定性与动态表现的需求相互冲突。本文提出一种多步态学习方法,使类人机器人通过一致的策略结构、动作空间和奖励函数掌握五种不同步态:步行、正步走、跑步、爬楼梯和跳跃。核心贡献是一种选择性对抗运动先验(Selective AMP)策略:对周期性且依赖稳定性的步态(步行、正步走、爬楼梯)应用AMP以加速收敛并抑制异常行为;而对高度动态的步态(跑步、跳跃)则刻意省略AMP,防止其过度约束运动自由度。策略采用PPO算法,在仿真中结合领域随机化进行训练,并通过零样本仿真到现实迁移部署于一台12-DOF类人机器人上。定量对比表明,选择性AMP在所有五种步态上均优于统一使用AMP的策略,实现了更快的收敛速度、更低的轨迹跟踪误差以及更高的稳定性相关步态成功率,同时未牺牲动态步态所需的灵活性。
原文摘要 · Abstract (English)
Learning diverse locomotion skills for humanoid robots in a unified reinforcement learning framework remains challenging due to the conflicting requirements of stability and dynamic expressiveness across different gaits. We present a multi-gait learning approach that enables a humanoid robot to master five distinct gaits -- walking, goose-stepping, running, stair climbing, and jumping -- using a consistent policy structure, action space, and reward formulation. The key contribution is a selective Adversarial Motion Prior (AMP) strategy: AMP is applied to periodic, stability-critical gaits (walking, goose-stepping, stair climbing) where it accelerates convergence and suppresses erratic behavior, while being deliberately omitted for highly dynamic gaits (running, jumping) where its regularization would over-constrain the motion. Policies are trained via PPO with domain randomization in simulation and deployed on a physical 12-DOF humanoid robot through zero-shot sim-to-real transfer. Quantitative comparisons demonstrate that selective AMP outperforms a uniform AMP policy across all five gaits, achieving faster convergence, lower tracking error, and higher success rates on stability-focused gaits without sacrificing the agility required for dynamic ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。