研究大步长梯度下降在过参数化回归中的收敛性,揭示了稳定性边缘的三种动态行为。
Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares
- 将过参数化解空间视为黎曼流形,分解梯度下降为切向与法向分量。
- 发现三类学习率下分别呈现有限时间收敛、幂律收敛和周期二轨道收敛。
- 为深度学习中大步长训练提供理论解释,适合关注优化机制的研究者。
经典优化理论保证小步长梯度下降(GD)时目标函数单调递减。然而,神经网络训练常采用大步长的‘稳定性边缘’策略,此时目标函数非单调下降,并表现出对平坦极小值的隐式偏好。本文在过参数化最小二乘设置下,量化分析大步长梯度下降的收敛速率。关键洞察在于:过参数化使全局最优解构成黎曼流形 $M$,可将GD动力学分解为沿 $M$ 的切向与垂直于 $M$ 的法向分量。切向分量对应目标函数尖锐度的黎曼梯度下降,法向分量则为分岔动力系统。据此,我们导出三类学习率下的收敛速率:(a) 亚临界区,瞬态不稳定性在有限时间内被克服,线性收敛至次优平坦的全局极小;(b) 临界区,不稳定性持续存在,以幂律收敛至最优平坦全局极小;(c) 超临界区,不稳定性持续存在,线性收敛至围绕最优平坦全局极小的周期二轨道。
原文摘要 · Abstract (English)
Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or ``stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the ``edge of stability", in which the objective decreases non-monotonically with an observed implicit bias towards flat minima. In this paper, we take a step toward quantifying this phenomenon by providing convergence rates for gradient descent with large learning rates in an overparametrised least squares setting. The key insight behind our analysis is that, as a consequence of overparametrisation, the set of global minimisers forms a Riemannian manifold $M$, which enables the decomposition of the GD dynamics into components parallel and orthogonal to $M$. The parallel component corresponds to Riemannian gradient descent on the objective sharpness, while the orthogonal component is a bifurcating dynamical system. This insight allows us to derive convergence rates in three regimes characterised by the learning rate size: (a) the subcritical regime, in which transient instability is overcome in finite time before linear convergence to a suboptimally flat global minimum; (b) the critical regime, in which instability persists for all time with a power-law convergence toward the optimally flat global minimum; and (c) the supercritical regime, in which instability persists for all time with linear convergence to an orbit of period two centred on the optimally flat global minimum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。