arXiv:2507.10030cs.RO2025-07被引 2

用进化策略优化强化学习控制,提升欠驱动机器人性能。

Finetuning Deep Reinforcement Learning Policies with Evolutionary Strategies for Control of Underactuated Robots

  • 先用SAC训练基础策略,再用自然进化策略精细调优。
  • 在IROS 2024真实机器人竞赛中显著提升控制表现。
  • 适合需要高鲁棒性与精准控制的复杂机器人任务。

深度强化学习(RL)已成为解决复杂控制问题的强大方法,尤其适用于欠驱动机器人系统。然而,某些情况下需进一步优化策略以实现最优性能与鲁棒性。本文提出一种基于进化策略(ES)的深度强化学习策略微调方法,用于提升欠驱动机器人的控制性能。首先使用软演员-评论家(SAC)算法,通过代理奖励函数近似复杂评分指标训练初始策略;随后采用分离式自然进化策略(SNES)进行零阶优化,直接优化原始得分。在IROS 2024 AI奥林匹克真实机器人挑战赛(RealAIGym)中的实验表明,该方法显著提升智能体表现,同时保持高鲁棒性,所获控制器在竞赛任务中达到竞争性得分,优于现有基线方法。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (RL) has emerged as a powerful method for addressing complex control problems, particularly those involving underactuated robotic systems. However, in some cases, policies may require refinement to achieve optimal performance and robustness aligned with specific task objectives. In this paper, we propose an approach for fine-tuning Deep RL policies using Evolutionary Strategies (ES) to enhance control performance for underactuated robots. Our method involves initially training an RL agent with Soft-Actor Critic (SAC) using a surrogate reward function designed to approximate complex specific scoring metrics. We subsequently refine this learned policy through a zero-order optimization step employing the Separable Natural Evolution Strategy (SNES), directly targeting the original score. Experimental evaluations conducted in the context of the 2nd AI Olympics with RealAIGym at IROS 2024 demonstrate that our evolutionary fine-tuning significantly improves agent performance while maintaining high robustness. The resulting controllers outperform established baselines, achieving competitive scores for the competition tasks.

强化学习进化策略机器人控制欠驱动系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。