提出可变学习率与动量的SGD高阶近似方法,填补理论空白。
First and Second Order Approximations to Stochastic Gradient Descent Methods with Momentum Terms
- 基于弱假设构建含时变参数的SGD一阶二阶近似
- 首次实现对时变学习率与动量的统一理论分析
- 适合研究优化算法收敛性与动态调参的学者
随机梯度下降(SGD)在优化问题中广泛应用。基于动量的改进方法在某些情况下表现更优,但多数依据来自经验而非严格证明。尽管可通过连续近似研究梯度下降的动态行为,现有工作仅限于固定学习率或无动量的SGD。本文在弱假设下,为允许学习率和动量参数随时间变化的SGD提供了近似结果,拓展了理论框架。
原文摘要 · Abstract (English)
Stochastic Gradient Descent (SGD) methods see many uses in optimization problems. Modifications to the algorithm, such as momentum-based SGD methods have been known to produce better results in certain cases. Much of this, however, is due to empirical information rather than rigorous proof. While the dynamics of gradient descent methods can be studied through continuous approximations, existing works only cover scenarios with constant learning rates or SGD without momentum terms. We present approximation results under weak assumptions for SGD that allow learning rates and momentum parameters to vary with respect to time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。