arXiv:2501.19104cs.LG2025-01ICML被引 5

揭示神经网络训练中类内差异消失的内在机制,解释为何模型越准越接近理想结构。

Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime

  • 从损失曲面角度分析梯度流如何自然导出类内一致结构
  • 证明低损失与小梯度点必近似满足类内无差异,且误差由残差损失控制
  • 适用于研究深度学习泛化能力的理论工作者

神经坍缩(Neural Collapse)是一种现象,即训练良好的神经网络最后一层表示趋于高度结构化的几何形态。本文聚焦其最基本性质——NC1:类内变异性趋近于零。现有理论多基于数据无关的无约束特征模型,而本文从数据特定视角出发,分析三层数值网络在均场范式下的行为,前两层处于均场极限,后接线性层。我们建立了一个核心联系:低经验损失与小梯度范数的点(即接近驻点)近似满足NC1,且逼近程度由残差损失和梯度范数决定。进一步证明:(i) 均方误差上的梯度流会收敛至具有小经验损失的NC1解;(ii) 对于充分分离的数据分布,NC1与测试损失趋零可同时达成。这与实验观察一致:训练过程中神经坍缩出现的同时,模型测试误差接近零。总体表明,NC1源于梯度训练所诱导的损失曲面特性,并揭示了特定数据分布下NC1与小测试误差的共现性。

原文摘要 · Abstract (English)

Neural Collapse is a phenomenon where the last-layer representations of a well-trained neural network converge to a highly structured geometry. In this paper, we focus on its first (and most basic) property, known as NC1: the within-class variability vanishes. While prior theoretical studies establish the occurrence of NC1 via the data-agnostic unconstrained features model, our work adopts a data-specific perspective, analyzing NC1 in a three-layer neural network, with the first two layers operating in the mean-field regime and followed by a linear layer. In particular, we establish a fundamental connection between NC1 and the loss landscape: we prove that points with small empirical loss and gradient norm (thus, close to being stationary) approximately satisfy NC1, and the closeness to NC1 is controlled by the residual loss and gradient norm. We then show that (i) gradient flow on the mean squared error converges to NC1 solutions with small empirical loss, and (ii) for well-separated data distributions, both NC1 and vanishing test loss are achieved simultaneously. This aligns with the empirical observation that NC1 emerges during training while models attain near-zero test error. Overall, our results demonstrate that NC1 arises from gradient training due to the properties of the loss landscape, and they show the co-occurrence of NC1 and small test error for certain data distributions.

神经坍缩泛化理论均场分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。