arXiv:2510.20905cs.LGmath.PR2025-10被引 2

通过重尾噪声控制,提升SGD避开尖锐极小值的能力

Global Dynamics of Heavy-Tailed SGDs in Nonconvex Loss Landscape: Characterization and Control

  • 引入重尾噪声并截断,改变SGD全局动态行为
  • 实验验证该方法使模型收敛到更平坦的极小值点
  • 适合关注模型泛化性能与优化机制的研究者

随机梯度下降(SGD)及其变体推动了现代人工智能发展,但其理论理解远落后于实际应用。普遍认为SGD能避免损失曲面中与较差泛化性相关的尖锐局部极小值。为揭示这一现象并进一步增强其能力,必须突破传统局部收敛分析,全面理解SGD的全局动力学。本文基于Wang和Rhee(2023)的大型偏差与亚稳态分析,构建了一套技术工具,精确刻画了重尾SGD的全局动态。特别地,我们发现深度学习中一个有趣现象:在训练过程中注入并随后截断重尾噪声,可几乎完全避免尖锐极小值,从而提升测试数据上的泛化性能。模拟与深度学习实验均证实,采用梯度裁剪的重尾SGD能找到几何更平坦的局部极小值,并实现更优泛化表现。

原文摘要 · Abstract (English)

Stochastic gradient descent (SGD) and its variants enable modern artificial intelligence. However, theoretical understanding lags far behind their empirical success. It is widely believed that SGD has a curious ability to avoid sharp local minima in the loss landscape, which are associated with poor generalization. To unravel this mystery and further enhance such capability of SGDs, it is imperative to go beyond the traditional local convergence analysis and obtain a comprehensive understanding of SGDs' global dynamics. In this paper, we develop a set of technical machinery based on the recent large deviations and metastability analysis in Wang and Rhee (2023) and obtain sharp characterization of the global dynamics of heavy-tailed SGDs. In particular, we reveal a fascinating phenomenon in deep learning: by injecting and then truncating heavy-tailed noises during the training phase, SGD can almost completely avoid sharp minima and achieve better generalization performance for the test data. Simulation and deep learning experiments confirm our theoretical prediction that heavy-tailed SGD with gradient clipping finds local minima with a more flat geometry and achieves better generalization performance.

优化算法泛化性重尾噪声SGD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。