arXiv:2505.12982cs.LGcs.NE2025-05被引 2

用深度强化学习优化遗传算法多参数控制,提速超27%。

Multi-parameter Control for the $(1+(λ,λ))$-GA on OneMax via Deep Reinforcement Learning

  • 用深度强化学习分别调控四个关键参数,实现动态优化
  • 新策略在4万规模问题上比最优已有策略快13%
  • 成果可直接用于改进经典遗传算法的运行效率

已知进化算法可通过动态调整关键参数来提升性能,适应优化过程不同阶段。以$(1+(λ,λ))$遗传算法优化OneMax函数为例,最优参数策略可实现线性期望运行时间,而静态参数无法达成。尽管已有研究关注单一参数调控,但多参数联合调控仍极具挑战。本文重新审视该问题,将四个核心参数解耦,采用前沿深度强化学习技术逼近高效控制策略。结果显示,虽训练困难,但成功后其性能显著优于所有现有策略。基于学习结果,我们提出一种简单控制策略,在所有测试规模(最高达40,000)下,相比理论推荐默认设置提速27%,较最强现有irace调优策略提速13%。

原文摘要 · Abstract (English)

It is well known that evolutionary algorithms can benefit from dynamic choices of the key parameters that control their behavior, to adjust their search strategy to the different stages of the optimization process. A prominent example where dynamic parameter choices have shown a provable super-constant speed-up is the $(1+(λ,λ))$ Genetic Algorithm optimizing the OneMax function. While optimal parameter control policies result in linear expected running times, this is not possible with static parameter choices. This result has spurred a lot of interest in parameter control policies. However, many works, in particular theoretical running time analyses, focus on controlling one single parameter. Deriving policies for controlling multiple parameters remains very challenging. In this work we reconsider the problem of the $(1+(λ,λ))$ Genetic Algorithm optimizing OneMax. We decouple its four main parameters and investigate how well state-of-the-art deep reinforcement learning techniques can approximate good control policies. We show that although making deep reinforcement learning learn effectively is a challenging task, once it works, it is very powerful and is able to find policies that outperform all previously known control policies on the same benchmark. Based on the results found through reinforcement learning, we derive a simple control policy that consistently outperforms the default theory-recommended setting by $27\%$ and the irace-tuned policy, the strongest existing control policy on this benchmark, by $13\%$, for all tested problem sizes up to $40{,}000$.

遗传算法强化学习参数控制优化加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。