arXiv:2411.08070cs.ROcs.NE2024-11被引 1

用多目标进化算法自动设计训练课程,提升四足机器人运动学习稳定性。

Multi-Objective Algorithms for Learning Open-Ended Robotic Problems

  • 将速度指令映射到目标空间,同时优化性能与多样性。
  • 在复杂场景下错误率比最优基线降低19%至44%。
  • 适合研究开放性机器人任务的算法与系统设计者。

四足运动是拓展自主车辆能力范围的关键开放性问题。传统强化学习常因训练不稳定和样本效率低而受限。本文提出一种新方法——多目标学习(MOL),利用多目标进化算法作为自动课程学习机制,通过将速度指令投影到目标空间,并同时优化性能与多样性,显著提升学习过程的稳定性与适应性。在MuJoCo物理仿真环境中测试表明,该方法在困难场景下相较于最优基线算法,统一评估下错误率减少19%,定制评估下减少44%。本工作为四足机器人训练提供了一个稳健框架,有望推动机器人运动及开放性机器人问题的显著进展。

原文摘要 · Abstract (English)

Quadrupedal locomotion is a complex, open-ended problem vital to expanding autonomous vehicle reach. Traditional reinforcement learning approaches often fall short due to training instability and sample inefficiency. We propose a novel method leveraging multi-objective evolutionary algorithms as an automatic curriculum learning mechanism, which we named Multi-Objective Learning (MOL). Our approach significantly enhances the learning process by projecting velocity commands into an objective space and optimizing for both performance and diversity. Tested within the MuJoCo physics simulator, our method demonstrates superior stability and adaptability compared to baseline approaches. As such, it achieved 19\% and 44\% fewer errors against our best baseline algorithm in difficult scenarios based on a uniform and tailored evaluation respectively. This work introduces a robust framework for training quadrupedal robots, promising significant advancements in robotic locomotion and open-ended robotic problems.

机器人运动多目标优化进化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。