揭示了随机优化中步长过大会导致性能下降的理论机制
Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
- 提出衡量步长敏感性的关键量,分析方法稳定性
- 证明自适应方法在凸问题下比SGD更优,且可量化优势
- 实验验证理论边界与实际性能变化趋势一致
我们从理论上分析了随机优化方法对步长的敏感性。识别出每种方法的关键量,该量描述了当步长过大时性能如何退化。对于凸问题,该量直接影响方法的次优性界。最重要的是,我们的分析为自适应步长方法(如SPS或NGN)比SGD更具鲁棒性提供了直接理论证据,使我们能够超越经验评估来量化这些方法的优势。最后,实验表明,即使在非凸问题上,我们的理论边界也能定性地反映实际性能随步长的变化趋势。
原文摘要 · Abstract (English)
We present a theoretical analysis of stochastic optimization methods in terms of their sensitivity with respect to the step size. We identify a key quantity that, for each method, describes how the performance degrades as the step size becomes too large. For convex problems, we show that this quantity directly impacts the suboptimality bound of the method. Most importantly, our analysis provides direct theoretical evidence that adaptive step-size methods, such as SPS or NGN, are more robust than SGD. This allows us to quantify the advantage of these adaptive methods beyond empirical evaluation. Finally, we show through experiments that our theoretical bound qualitatively mirrors the actual performance as a function of the step size, even for non-convex problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。