揭示了SGD与加速SGD在高维二次优化中的最优边界
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
- 分析带衰减学习率和动量的SGD收敛性
- 证明在幂律衰减问题上可达最优收敛速率
- 首次明确SGD与加速SGD的适用边界
随机梯度下降(SGD)是机器学习中广泛使用的算法,尤其适用于神经网络训练。近期研究显示,在高维设定下,针对标准二次优化或线性回归问题,SGD可实现良好泛化性能。然而,一个基本问题尚未充分探讨:在哪些高维学习问题中,SGD及其加速变体能够达到最优?本文研究了实践中两个关键成分:指数衰减学习率和动量。我们建立了带动量的加速SGD(ASGD)的收敛上界,并提出了具体的学习问题类别,使得SGD或ASGD能达到最小最大收敛速率。目标函数的刻画基于(函数)线性回归中的标准幂律衰减。结果揭示了关于SGD学习偏见的新见解:(i) SGD在权重受无穷范数约束的“密集”特征学习中高效;(ii) 对于无饱和效应的简单问题,SGD表现良好;(iii) 当学习问题较难时,动量可使收敛速率提升一个阶。据我们所知,这是首个在温和设定下清晰识别出SGD与ASGD最优边界的论文。
原文摘要 · Abstract (English)
Stochastic gradient descent (SGD) is a widely used algorithm in machine learning, particularly for neural network training. Recent studies on SGD for canonical quadratic optimization or linear regression show it attains well generalization under suitable high-dimensional settings. However, a fundamental question -- for what kinds of high-dimensional learning problems SGD and its accelerated variants can achieve optimality has yet to be well studied. This paper investigates SGD with two essential components in practice: exponentially decaying step size schedule and momentum. We establish the convergence upper bound for momentum accelerated SGD (ASGD) and propose concrete classes of learning problems under which SGD or ASGD achieves min-max optimal convergence rates. The characterization of the target function is based on standard power-law decays in (functional) linear regression. Our results unveil new insights for understanding the learning bias of SGD: (i) SGD is efficient in learning ``dense'' features where the corresponding weights are subject to an infinity norm constraint; (ii) SGD is efficient for easy problem without suffering from the saturation effect; (iii) momentum can accelerate the convergence rate by order when the learning problem is relatively hard. To our knowledge, this is the first work to clearly identify the optimal boundary of SGD versus ASGD for the problem under mild settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。