通过熵控内在动机,提升四足机器人在复杂地形的运动稳定性与能效。
Entropy-Controlled Intrinsic Motivation Reinforcement Learning for Quadruped Robot Locomotion in Complex Terrains
- 引入熵控内在动机机制,动态调节探索强度以避免过早收敛。
- 在六类地形上任务奖励提升4%~12%,关节加速度下降20%~32%。
- 适合需要高稳定性和低能耗的复杂环境机器人控制任务。
学习是生物与人工系统模仿智能行为的基础。尽管经典PPO系列深度强化学习算法因稳定性和样本效率被广泛用于四足机器人步态训练,但其在实验与仿真中常出现过早收敛,导致运动性能不佳。为此,本文提出熵控内在动机(ECIM)算法,通过结合内在动机与自适应探索,缓解过早收敛问题。实验在Isaac Gym平台进行,覆盖六类地形:上坡、下坡、崎岖不平地面、上楼梯、下楼梯及平坦地面。相比基线方法,本方法任务奖励提升4%~12%,峰值躯体俯仰振荡降低23%~29%,关节加速度下降20%~32%,关节扭矩消耗减少11%~20%。结果表明,ECIM通过熵控与内在动机协同,显著提升四足机器人在复杂地形下的运动稳定性与能效,适用于实际复杂机器人控制任务。
原文摘要 · Abstract (English)
Learning is the basis of both biological and artificial systems when it comes to mimicking intelligent behaviors. From the classical PPO (Proximal Policy Optimization), there is a series of deep reinforcement learning algorithms which are widely used in training locomotion policies for quadrupedal robots because of their stability and sample efficiency. However, among all these variants, experiments and simulations often converge prematurely, leading to suboptimal locomotion and reduced task performance. Therefore, in this paper, we introduce Entropy-Controlled Intrinsic Motivation (ECIM), an entropy-based reinforcement learning algorithm in contrast with the PPO series, that can reduce premature convergence by combining intrinsic motivation with adaptive exploration. For experiments, in order to parallel with other baselines, we chose to apply it in Isaac Gym across six terrain categories: upward slopes, downward slopes, uneven rough terrain, ascending stairs, descending stairs, and flat ground as widely used. For comparison, our experiments consistently achieve better performance: task rewards increase by 4--12%, peak body pitch oscillation is reduced by 23--29%, joint acceleration decreases by 20--32%, and joint torque consumption declines by 11--20%. Overall, our model ECIM, by combining entropy control and intrinsic motivation control, achieves better results in stability across different terrains for quadrupedal locomotion, and at the same time reduces energetic cost and makes it a practical choice for complex robotic control tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。