用自适应批量估计提升无梯度优化的精度与效率
Enhanced Derivative-Free Optimization Using Adaptive Correlation-Induced Finite Difference Estimators
- 基于相关性诱导的批量有限差分法,提高梯度估计精度
- 动态调整每轮采样数,实现样本高效利用
- 无需减小步长也能收敛,适合实际工程优化
基于梯度的方法在无梯度优化(DFO)中表现良好,通常使用有限差分(FD)估计作为梯度替代。传统随机逼近方法如Kiefer-Wolfowitz(KW)和同时扰动随机逼近(SPSA)每轮仅用两个样本,导致梯度估计不精确,且需递减步长以保证收敛。本文提出一种高效的关联诱导有限差分估计,为基于批次的估计方法。进一步设计自适应采样策略,动态确定每轮迭代的批次大小。结合两者,构建新算法,在提升梯度估计效率与样本效率的同时,确保一致性。尽管每轮使用一批样本,该算法仍保持与KW和SPSA相同的收敛速率。此外,提出一种新型随机线搜索技术,用于实践中自适应调节步长。大量数值实验验证了所提算法的优越性。
原文摘要 · Abstract (English)
Gradient-based methods are well-suited for derivative-free optimization (DFO), where finite-difference (FD) estimates are commonly used as gradient surrogates. Traditional stochastic approximation methods, such as Kiefer-Wolfowitz (KW) and simultaneous perturbation stochastic approximation (SPSA), typically utilize only two samples per iteration, resulting in imprecise gradient estimates and necessitating diminishing step sizes for convergence. In this paper, we first explore an efficient FD estimate, referred to as correlation-induced FD estimate, which is a batch-based estimate. Then, we propose an adaptive sampling strategy that dynamically determines the batch size at each iteration. By combining these two components, we develop an algorithm designed to enhance DFO in terms of both gradient estimation efficiency and sample efficiency. Furthermore, we establish the consistency of our proposed algorithm and demonstrate that, despite using a batch of samples per iteration, it achieves the same convergence rate as the KW and SPSA methods. Additionally, we propose a novel stochastic line search technique to adaptively tune the step size in practice. Finally, comprehensive numerical experiments confirm the superior empirical performance of the proposed algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。