揭示SGD为何偏好泛化能力强的平坦解
Noise-Driven Exploration and Transient Freezing Select Flat Minima in Stochastic Gradient Descent
- 通过噪声驱动探索与瞬时冻结机制解释优化路径
- 更强噪声延缓冻结,提升收敛到平坦极小值概率
- 为设计更优优化算法提供物理层面指导
随机梯度下降(SGD)是深度学习的核心,但其为何偏好更平坦、泛化性更好的解仍不明确。本文分析训练动态,发现一种非平衡机制:训练初期,SGD轨迹反复逃离尖锐山谷并迁移至损失曲面更平坦区域,随后被限制在最终盆地。通过可解析的物理模型,我们表明SGD噪声重塑损失曲面为有效势能场,优先稳定平坦解。进一步发现瞬时冻结机制:随着训练进行,平坦化势场抑制不同山谷间的跃迁。更强的噪声延迟冻结过程,延长探索期,从而提高收敛至平坦极小值的概率。这些结果构建了学习动态、损失曲面几何与泛化性能之间的统一物理框架,并为优化算法设计提供新原则。
原文摘要 · Abstract (English)
Stochastic gradient descent (SGD) is central to deep learning, yet the dynamical origin of its preference for flatter, more generalizable solutions remains unclear. Here, by analyzing SGD learning dynamics, we identify a nonequilibrium mechanism that governs solution selection during training. Numerical experiments reveal a transient exploratory phase in which SGD trajectories repeatedly escape sharp valleys and migrate toward flatter regions of the loss landscape before becoming confined to a final basin. Using a tractable physical model, we show that SGD noise reshapes the loss landscape into an effective potential that preferentially stabilizes flat solutions. We further uncover a transient freezing mechanism: as training progresses, the flattening landscape suppresses transitions between competing valleys. Stronger SGD noise delays this freezing transition, prolonging the exploratory phase and thereby increasing the probability of convergence to flatter minima. Together, these results provide a unified physical framework connecting learning dynamics, loss-landscape geometry, and generalization, and suggest guiding principles for the design of more effective optimization algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。