arXiv:2502.20265cs.LGcs.NE2025-02被引 5

奖励设计不当会阻碍强化学习配置遗传算法,改进后显著提升学习效果。

On the Importance of Reward Design in Reinforcement Learning-based Dynamic Algorithm Configuration: A Case Study on OneMax with (1+($λ$,$λ$))-GA

  • 用奖励塑形增强智能体探索能力,改善动态配置
  • 差的奖励导致学习发散,无法适应不同规模问题
  • 适合研究强化学习与进化算法结合的学者参考

动态算法配置(DAC)近年来受到广泛关注,尤其在机器学习与深度学习背景下。众多研究利用强化学习(RL)在决策上的鲁棒性来应对算法配置的优化挑战。然而,让RL智能体有效工作并非易事,尤其是奖励设计,往往需要大量依赖领域知识的手工设计。本文通过控制(1+(λ,λ))-GA在OneMax问题中种群大小的案例研究,探讨了奖励设计的重要性。我们发现,不良的奖励设计会因缺乏探索而阻碍智能体学习最优策略,导致可扩展性差和学习发散问题。为此,我们提出采用奖励塑形机制,以促进智能体对环境的更优探索。本工作不仅展示了RL在动态配置(1+(λ,λ))-GA中的能力,还验证了奖励塑形在提升智能体对不同规模OneMax问题可扩展性方面的优势。

原文摘要 · Abstract (English)

Dynamic Algorithm Configuration (DAC) has garnered significant attention in recent years, particularly in the prevalence of machine learning and deep learning algorithms. Numerous studies have leveraged the robustness of decision-making in Reinforcement Learning (RL) to address the optimization challenges associated with algorithm configuration. However, making an RL agent work properly is a non-trivial task, especially in reward design, which necessitates a substantial amount of handcrafted knowledge based on domain expertise. In this work, we study the importance of reward design in the context of DAC via a case study on controlling the population size of the $(1+(λ,λ))$-GA optimizing OneMax. We observed that a poorly designed reward can hinder the RL agent's ability to learn an optimal policy because of a lack of exploration, leading to both scalability and learning divergence issues. To address those challenges, we propose the application of a reward shaping mechanism to facilitate enhanced exploration of the environment by the RL agent. Our work not only demonstrates the ability of RL in dynamically configuring the $(1+(λ,λ))$-GA, but also confirms the advantages of reward shaping in the scalability of RL agents across various sizes of OneMax problems.

强化学习算法配置奖励设计进化计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。