arXiv:2410.16401cs.LGmath.ST2024-10ICML被引 4

小步长下标签噪声SGD让两层网络收敛到秩一线性特征,解释了模型为何简单。

Simplicity Bias via Global Convergence of Sharpness Minimization

  • 通过最小化损失曲面尖锐度,发现标签噪声SGD会收敛到低秩结构。
  • 在高维数据与特定激活下,收敛后所有神经元共享同一线性特征。
  • 揭示尖锐度与模型几何的关联,适合研究泛化机制的学者参考。

神经网络卓越的泛化能力通常归因于SGD的隐式偏差,常产生更简单的(如线性)和低秩特征。近期研究提供了实证与理论证据,表明特定变体的SGD(如标签噪声SGD)倾向于收敛到损失曲面中平坦区域。尽管平坦解常被认为“简单”,但其与最终模型简洁性(如低秩)的联系尚未清晰。本文研究一类两层神经网络中尖锐度最小化所导致的简洁结构。我们证明:在任意高维训练数据和某些激活函数下,当步长足够小时,标签噪声SGD总收敛至一个所有神经元共享单一线性特征的网络,即实现秩一特征矩阵。关键技术贡献在于:标签噪声SGD始终在零损失模型流形上最小化尖锐度。过程中发现一种新性质——近似驻点处损失迹的海森矩阵具有局部测地凸性,将尖锐度与流形几何相联。该工具可能具独立研究价值。

原文摘要 · Abstract (English)

The remarkable generalization ability of neural networks is usually attributed to the implicit bias of SGD, which often yields models with lower complexity using simpler (e.g. linear) and low-rank features. Recent works have provided empirical and theoretical evidence for the bias of particular variants of SGD (such as label noise SGD) toward flatter regions of the loss landscape. Despite the folklore intuition that flat solutions are 'simple', the connection with the simplicity of the final trained model (e.g. low-rank) is not well understood. In this work, we take a step toward bridging this gap by studying the simplicity structure that arises from minimizers of the sharpness for a class of two-layer neural networks. We show that, for any high dimensional training data and certain activations, with small enough step size, label noise SGD always converges to a network that replicates a single linear feature across all neurons; thereby, implying a simple rank one feature matrix. To obtain this result, our main technical contribution is to show that label noise SGD always minimizes the sharpness on the manifold of models with zero loss for two-layer networks. Along the way, we discover a novel property -- a local geodesic convexity -- of the trace of Hessian of the loss at approximate stationary points on the manifold of zero loss, which links sharpness to the geometry of the manifold. This tool may be of independent interest.

泛化能力神经网络优化理论低秩结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。