arXiv:2604.07224cs.RO2026-04中稿 · the 11th Internati…

用进化强化学习提升四足机器人在复杂地形的适应能力

Robust Quadruped Locomotion via Evolutionary Reinforcement Learning

  • 结合梯度优化与种群探索,增强策略鲁棒性
  • 进化方法在粗糙地形上表现远超传统强化学习(最高19574.33分)
  • 适合关注实际部署中泛化能力的研究者

深度强化学习在四足机器人行走任务中表现优异,但仿真训练的策略在环境变化时难以迁移。本文评估了四种方法在模拟行走任务中的表现:DDPG、TD3以及两种基于交叉熵的变体CEM-DDPG和CEM-TD3。所有智能体均在平坦地形上训练,随后在训练未涉及的粗糙地形上测试。在平坦地形上,TD3表现最佳,平均奖励为5927.26;而CEM-TD3在训练和评估中均获得最高奖励,达17611.41。在粗糙地形迁移测试中,传统强化学习方法性能大幅下降:DDPG得分为-1016.32,TD3为-99.73;而进化方法仍保持较强能力,其中CEM-TD3在粗糙地形上的平均奖励达到19574.33。结果表明,引入进化搜索可减少过拟合,显著提升运动策略在不同环境下的鲁棒性。

原文摘要 · Abstract (English)

Deep reinforcement learning has recently achieved strong results in quadrupedal locomotion, yet policies trained in simulation often fail to transfer when the environment changes. Evolutionary reinforcement learning aims to address this limitation by combining gradient-based policy optimisation with population-driven exploration. This work evaluates four methods on a simulated walking task: DDPG, TD3, and two Cross-Entropy-based variants CEM-DDPG and CEM-TD3. All agents are trained on flat terrain and later tested both on this domain and on a rough terrain not encountered during training. TD3 performs best among the standard deep RL baselines on flat ground with a mean reward of 5927.26, while CEM-TD3 achieves the highest rewards overall during training and evaluation 17611.41. Under the rough-terrain transfer test, performance of the deep RL methods drops sharply. DDPG achieves -1016.32 and TD3 achieves -99.73, whereas the evolutionary variants retain much of their capability. CEM-TD3 records the strongest transfer performance with a mean reward of 19574.33. These findings suggest that incorporating evolutionary search can reduce overfitting and improve policy robustness in locomotion tasks, particularly when deployment conditions differ from those seen during training.

强化学习四足机器人迁移能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。