arXiv:2602.00532cs.NEcs.LG2026-02

用强化学习自动调节约束松弛,提升复杂优化效率。

Reinforcement Learning-assisted Constraint Relaxation for Constrained Expensive Optimization

  • 通过强化学习动态调整约束松弛程度,适应优化过程变化。
  • 在CEC 2017基准上表现优于或媲美竞赛优胜方法。
  • 适合高成本、依赖专家设计的复杂优化场景使用。

约束处理在解决现实复杂优化问题中至关重要。尽管过去数十年研究不断,现有方法仍主要依赖人工设计,通用性不足。受元黑箱优化中自动化算法设计进展启发,本文提出一种基于强化学习的自适应、通用约束处理策略。我们构建了定制化的马尔可夫决策过程,利用深度Q网络根据优化动态特征控制约束松弛水平,实现目标探索与可行域探索之间的灵活权衡,从而提升优化性能。模型在受限评估预算(昂贵情形)下的CEC 2017约束优化基准上训练,并与近期CEC/GECCO竞赛优胜者等强基线对比。大量实验表明,无论采用留一交叉验证还是普通训练-测试划分,本方法均表现竞争力甚至超越基线。进一步分析与消融实验揭示了设计的关键洞见。

原文摘要 · Abstract (English)

Constraint handling plays a key role in solving realistic complex optimization problems. Though intensively discussed in the last few decades, existing constraint handling techniques predominantly rely on human experts' designs, which more or less fall short in utility towards general cases. Motivated by recent progress in Meta-Black-Box Optimization where automated algorithm design can be learned to boost optimization performance, in this paper, we propose learning effective, adaptive and generalizable constraint handling policy through reinforcement learning. Specifically, a tailored Markov Decision Process is first formulated, where given optimization dynamics features, a deep Q-network-based policy controls the constraint relaxation level along the underlying optimization process. Such adaptive constraint handling provides flexible tradeoff between objective-oriented exploitation and feasible-region-oriented exploration, and hence leads to promising optimization performance. We train our approach on CEC 2017 Constrained Optimization benchmark with limited evaluation budget condition (expensive cases) and compare the trained constraint handling policy to strong baselines such as recent winners in CEC/GECCO competitions. Extensive experimental results show that our approach performs competitively or even surpasses the compared baselines under either Leave-one-out cross-validation or ordinary train-test split validation. Further analysis and ablation studies reveal key insights in our designs.

强化学习优化约束处理元优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。