arXiv:2608.02073math.OCcs.LG2026-08

利用实际输入信息加速进化策略,提升优化效率。

Accelerating Evolutionary Strategy via Rao-Blackwellizing Realization of Uncertain Input

  • 通过罗-布莱克韦尔化降低梯度估计方差,利用可观测的实际输入信息。
  • 在连续优化与强化学习任务中,收敛速度显著快于传统进化策略。
  • 适合有输入不确定性的制造、控制与强化学习场景,尤其注重效率的部署。

我们研究输入不确定性下的优化(OIU),即目标函数的输入而非函数本身存在不确定性。这类问题出现在具有生产公差的制造过程、存在执行噪声的物理系统控制、专家混合模型以及强化学习中。现有方法通常仅使用目标函数值,忽略了可观测的实际输入信息。本文提出:被丢弃的实际输入信息是否有助于加速优化?针对进化策略(ES),我们从理论上证明,利用实际输入信息可通过罗-布莱克韦尔化降低梯度估计方差。基于此,我们提出表型加速进化策略(PAES),作为应对输入不确定性的ES改进方法。数值实验表明,从简单连续优化到强化学习基准测试,PAES均比标准进化策略收敛更快。

原文摘要 · Abstract (English)

We investigate Optimization under Input Uncertainty (OIU), in which the input to the objective function, rather than the objective function itself, is subject to uncertainty. OIU appears in manufacturing processes with production tolerance, control of physical systems with actuation noise, Mixture of Experts, and Reinforcement Learning (RL). Most of the existing approaches solve OIU by using the value of the objective function but discard the information of the realized input, even though the realized input is observable in various applications. The question here is whether the discarded information of the realized input is useful to accelerate the optimization process. We affirmatively answer this question for Evolutionary Strategy (ES) by theoretically showing that the information of the realized input can reduce the variance of the gradient estimator via Rao-Blackwellization. Using the Rao-Blackwellized gradient estimator, we propose Phenotype-Accelerated Evolutionary Strategy (PAES), which is a refinement of ES for OIU. Numerical experiments show that PAES converges faster than the usual ES from simple continuous optimization problems to RL benchmarks.

进化策略不确定性优化强化学习加速优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。