arXiv:2606.30930stat.MLcs.LG2026-06

SGD在大学习率下能自动稳定训练,突破传统理论限制。

SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates

  • 利用随机性使梯度下降在稳定与波动间交替,实现自我调节。
  • 即使学习率极大,仍能保证最佳迭代点收敛,损失可控下降。
  • 适合研究大步长训练机制或设计鲁棒优化算法的读者。

现代深度学习常以远超经典优化理论允许的学习率运行,处于稳定性边缘。以往研究多聚焦于确定性梯度下降,对随机情形分析不足。本文首次为多分类交叉熵损失下的随机梯度下降(SGD)提供精确收敛保证,涵盖线性分类器和两层神经网络。研究表明,SGD的随机性导致系统在由曲率主导的波动状态与损失可控下降的稳定状态间交替。尽管如此,我们证明了SGD具备自稳定能力:在固定迭代次数内使轨迹回归稳定,确保最佳迭代点收敛,即使在大学习率下依然有效。实验验证了理论结果,并展示了大步长下SGD的优势。

原文摘要 · Abstract (English)

Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory. Most prior analyses of the edge of stability phenomenon focus on deterministic gradient descent, leaving the stochastic setting largely unexplored. In this work, we provide sharp convergence guarantees for Stochastic Gradient Descent (SGD) applied to the multiclass cross-entropy loss, for both linear classifiers and two-layer neural networks. We show that the stochasticity of SGD may cause the dynamics to alternate between an edge-of-stability regime that is dominated by curvature-driven oscillations, and a stable regime in which the expected loss decreases at a controlled rate. Despite that, we prove that SGD self-stabilizes the dynamics, ensuring that the iterates return to stability in a fixed number of iterations and allowing convergence in the best-iterate sense even with large learning rates. Experiments validate our theoretical findings and illustrate the benefits of SGD in the large-stepsize regime.

随机优化大步长训练自稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。