为深度学习中的动量优化提供理论保障,突破传统凸优化限制。
Nesterov acceleration in benignly non-convex landscapes
- 在良性非凸环境下分析Nesterov加速梯度法,放宽几何假设。
- 证明在过参数化深度学习中,局部优化可获得与凸情况相当的收敛保证。
- 适用于连续/离散版本NAG,以及带加性或乘加性噪声的随机梯度场景。
尽管动量类优化算法广泛应用于深度学习中典型的非凸优化问题,其理论分析长期局限于凸或强凸设定。本文部分弥合了理论与实践之间的差距,证明在具有‘良性’非凸性的优化问题中,几乎相同的收敛保证仍可成立。我们表明,这类较弱的几何假设在过参数化深度学习中至少在局部是合理的。该结果进一步推广至Nesterov加速梯度下降(NAG)的连续时间模型、经典的离散时间版本,以及带有纯加性噪声和同时具备加性和乘性缩放特性的噪声的随机梯度版本。
原文摘要 · Abstract (English)
While momentum-based optimization algorithms are commonly used in the notoriously non-convex optimization problems of deep learning, their analysis has historically been restricted to the convex and strongly convex setting. In this article, we partially close this gap between theory and practice and demonstrate that virtually identical guarantees can be obtained in optimization problems with a `benign' non-convexity. We show that these weaker geometric assumptions are well justified in overparametrized deep learning, at least locally. Variations of this result are obtained for a continuous time model of Nesterov's accelerated gradient descent algorithm (NAG), the classical discrete time version of NAG, and versions of NAG with stochastic gradient estimates with purely additive noise and with noise that exhibits both additive and multiplicative scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。