arXiv:2606.06555cs.NEcs.LG2026-06中稿 · ICML

在有限评估预算下,用概率方法提升进化策略的搜索深度。

Depth over Fidelity in Fixed-Budget Noisy Evolution Strategies

论文配图:Depth over Fidelity in Fixed-Budget Noisy Evolution Strategies
图 1 · 摘自论文原文
  • 用条件期望排名权重替代硬性排序,降低噪声干扰
  • 在高噪声、低预算场景中显著提升优化性能
  • 适合超参数调优与强化学习策略搜索等任务

固定评估预算下的噪声进化策略面临深度与精度的权衡:用于降噪的评估会减少分布更新次数。本文主张优先考虑深度,并提出概率精英成员(PEM),将进化策略中的硬性排名权重替换为基于排名不确定性的条件期望权重,从而在保持条件均值更新的同时降低更新方差,实现对噪声排名步长的Rao-Blackwell化。通过残差自助法(RB-PEM)实现PEM,每代开销受控,并引入自适应探查-切换机制应对低噪声情形。在COCO bbob-noisy基准及外部任务(包括强化学习策略搜索和超参数优化)中,RB-PEM在高误排序、预算受限条件下表现稳定提升。

原文摘要 · Abstract (English)

Noisy evolution strategies under fixed evaluation budgets face a depth-fidelity trade-off: spending evaluations to denoise intra-generation rankings reduces the number of distribution updates the optimizer can execute. We argue for depth over fidelity and propose probabilistic elite membership (PEM), which replaces hard rank-based weights in evolution strategies with conditional expected rank weights that integrate over ranking uncertainty. PEM preserves the conditional mean update while reducing conditional update dispersion, a Rao-Blackwellization of the noisy rank-based step. We instantiate PEM via residual bootstrapping (RB-PEM) with capped per-generation overhead, complemented by an adaptive probe-and-switch mechanism for low-noise regimes. Across the COCO bbob-noisy suite and external tasks including RL policy search and hyperparameter optimization, RB-PEM achieves consistent gains in high-misranking, budget-constrained settings.

进化策略噪声优化超参数调优强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。