模仿动物运动演化,让四足机器人自适应复杂地形
Behavior evolution-inspired approach to walking gait reinforcement training for quadruped robots
- 引入参考步态自进化机制,结合遗传算法优化初始步态
- 在多种地形和模型参数下,步态适应性显著优于传统方法
- 适合研究机器人步态学习与自适应控制的科研人员
强化学习在四足机器人步态生成中表现优异,主要得益于其随机探索能力有助于实现自主步态。然而,尽管采用增量强化学习以利用肢体运动连续性提升训练成功率和运动平滑性,面对不同地形和外部扰动时仍存在适应性挑战。受动物运动行为演化的启发,本文提出一种参考步态的自改进机制,实现动作增量学习与参考动作自我优化的协同进化。进一步构建了四足步态强化训练新框架:首先使用遗传算法对任意足部轨迹的初始值进行全局概率搜索,以更优适应度更新参考轨迹;随后基于优化后的参考步态执行增量强化学习。该过程反复交替进行,最终训练出稳健的步态策略。通过仿真分析多种地形、模型尺寸及运动条件,结果表明该框架在地形适应性方面显著优于常规增量强化学习。
原文摘要 · Abstract (English)
Reinforcement learning method is extremely competitive in gait generation techniques for quadrupedal robot, which is mainly due to the fact that stochastic exploration in reinforcement training is beneficial to achieve an autonomous gait. Nevertheless, although incremental reinforcement learning is employed to improve training success and movement smoothness by relying on the continuity inherent during limb movements, challenges remain in adapting gait policy to diverse terrain and external disturbance. Inspired by the association between reinforcement learning and the evolution of animal motion behavior, a self-improvement mechanism for reference gait is introduced in this paper to enable incremental learning of action and self-improvement of reference action together to imitate the evolution of animal motion behavior. Further, a new framework for reinforcement training of quadruped gait is proposed. In this framework, genetic algorithm is specifically adopted to perform global probabilistic search for the initial value of the arbitrary foot trajectory to update the reference trajectory with better fitness. Subsequently, the improved reference gait is used for incremental reinforcement learning of gait. The above process is repeatedly and alternatively executed to finally train the gait policy. The analysis considering terrain, model dimensions, and locomotion condition is presented in detail based on simulation, and the results show that the framework is significantly more adaptive to terrain compared to regular incremental reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。