arXiv:2507.11274cs.LGmath.OC2025-07NeurIPS被引 16

提出SGD在平滑插值情形下的快速末次迭代收敛新理论。

Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime

  • 基于平滑凸损失与常数步长,推导末次迭代期望过失风险上界。
  • 当噪声为零时,收敛率达$O(1/\ oot\of{T})$,优于已有$O(T^{-1/4})$结果。
  • 适用于过参数化模型训练、持续学习遗忘分析等场景。

研究了在插值情形下(最优解处噪声为零或接近零)平滑凸目标函数的随机梯度下降(SGD)的总体收敛性。针对大(常数)步长下SGD末次迭代的行为,我们证明:在$T$步内,对$β$-光滑凸损失函数,步长$0 < η < 2/β$时,末次迭代的期望过剩风险为$\widetilde{O}(\frac{1}{η(2-βη) T^{1-βη/2}} + \frac{η}{(2-βη)^2} T^{βη/2} σ_\star^2)$,其中$σ_\star^2$为最优解处随机梯度方差。特别地,对优化步长可得近似最优$\widetilde{O}(1/T + σ_\star/\sqrt{T})$率,推广了Varre等(2021)在最小二乘回归外的结果;当$σ_\star=0$时,取$η=1/β$,收敛率可达$O(1/\sqrt{T})$,优于最近Evron等(2025)在可实现线性回归情形下$O(T^{-1/4})$的最佳已知结果。

原文摘要 · Abstract (English)

We study population convergence guarantees of stochastic gradient descent (SGD) for smooth convex objectives in the interpolation regime, where the noise at optimum is zero or near zero. The behavior of the last iterate of SGD in this setting -- particularly with large (constant) stepsizes -- has received growing attention in recent years due to implications for the training of over-parameterized models, as well as to analyzing forgetting in continual learning and to understanding the convergence of the randomized Kaczmarz method for solving linear systems. We establish that after $T$ steps of SGD on $β$-smooth convex loss functions with stepsize $0 < η< 2/β$, the last iterate exhibits expected excess risk $\widetilde{O}(\frac{1}{η(2-βη) T^{1-βη/2}} + \fracη{(2-βη)^2} T^{βη/2} σ_\star^2)$, where $σ_\star^2$ denotes the variance of the stochastic gradients at the optimum. In particular, for a well-tuned stepsize we obtain a near optimal $\widetilde{O}(1/T + σ_\star/\sqrt{T})$ rate for the last iterate, extending the results of Varre et al. (2021) beyond least squares regression; and when $σ_\star=0$ we obtain a rate of $\smash{O(1/\sqrt T)}$ with $η=1/β$, improving upon the best-known $\smash{O(T^{-1/4})}$ rate recently established by Evron et al. (2025) in the special case of realizable linear regression.

SGD收敛分析平滑优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。