揭示小批量噪声如何通过主导子空间波动降低模型尖锐度
Mini-batch Noise Lowers Sharpness via Dominant-Subspace Fluctuations
- 发现梯度在损失海塞矩阵主子空间的波动主导了尖锐度变化
- 证明小批量噪声在主子空间内产生修正项,使梯度下降更接近SGD的尖锐度演化
- 适用于研究优化器动态、模型泛化与尖锐度控制的学者
在SGD训练过程中,梯度常强烈对齐于损失函数海塞矩阵前k个特征向量张成的主导子空间。尽管这看似意味着损失减少主要发生在此空间,但先前研究表明该子空间内的更新对降低损失无实质贡献。本文提出,主导子空间不应视为损失减少的主要空间,而应理解为解释小批量SGD尖锐度动态的关键子空间。我们证明,主导方向上梯度波动的平均值会产生一个尖锐度修正项,并推导出由小批量噪声在主导方向上引起的尖锐度修正项。实验表明,将该修正项加入梯度下降(GD)后,GD的尖锐度演化轨迹更接近于SGD。
原文摘要 · Abstract (English)
During SGD training, the gradients often align strongly with the dominant subspace spanned by the top-$k$ eigenvectors of the Hessian of the loss. While this seems to naturally imply that loss reduction mainly occurs within this space, prior work has shown that updates within this dominant subspace make no meaningful progress in reducing the loss. In this work, we argue that the dominant subspace is better understood not as the main space for loss reduction, but as a key subspace for explaining the sharpness dynamics of mini-batch SGD. To explain the role of the dominant subspace in reducing top-$k$ sharpness, we show how the averaged gradient over fluctuations in the dominant directions produces a sharpness correction term, and derive a sharpness correction term induced by mini-batch noise in the dominant directions. Experimental results show that adding the derived correction term to GD brings the sharpness evolution of GD closer to that of SGD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。