用示范引导强化学习,让四足机器人学会稳定跳跃。
Explosive Jumping with Rigid and Articulated Soft Quadrupeds via Example Guided Reinforcement Learning
- 通过模仿优化轨迹生成的示范,逐步训练跳跃策略。
- 单次示范即学会多种距离跳跃,落地误差降低11.1%。
- 适合追求高鲁棒性跳跃的仿生机器人研究者。
实现四足机器人受控跳跃极具挑战性,尤其在引入被动柔顺设计时。本文采用基于模仿的深度强化学习与渐进式训练流程解决该问题。首先,通过模型基轨迹优化生成粗略跳跃示范,模仿学习跳跃技能;随后,将策略泛化至不同前后及侧向跳跃距离,并提升对未知地面不平的鲁棒性。此外,在极少调整奖励函数的前提下,成功训练出具有并联弹性结构的四足机器人跳跃策略。实验表明:(i)仅需单一示范即可学习多样化跳跃行为;(ii)相比刚性结构,具备并联弹性结构的机器人落地误差减少11.1%,能耗降低15.2%,峰值扭矩下降15.8%;(iii)仅依赖本体感知,即可在最大4cm高度扰动下实现可变距离跳跃,具备强鲁棒性。
原文摘要 · Abstract (English)
Achieving controlled jumping behaviour for a quadruped robot is a challenging task, especially when introducing passive compliance in mechanical design. This study addresses this challenge via imitation-based deep reinforcement learning with a progressive training process. To start, we learn the jumping skill by mimicking a coarse jumping example generated by model-based trajectory optimization. Subsequently, we generalize the learned policy to broader situations, including various distances in both forward and lateral directions, and then pursue robust jumping in unknown ground unevenness. In addition, without tuning the reward much, we learn the jumping policy for a quadruped with parallel elasticity. Results show that using the proposed method, i) the robot learns versatile jumps by learning only from a single demonstration, ii) the robot with parallel compliance reduces the landing error by 11.1%, saves energy cost by 15.2% and reduces the peak torque by 15.8%, compared to the rigid robot without parallel elasticity, iii) the robot can perform jumps of variable distances with robustness against ground unevenness (maximal 4cm height perturbations) using only proprioceptive perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。