arXiv:2505.15239cs.LGcs.AI2025-05NeurIPS被引 11

证明深度正则化模型在训练中会趋于神经坍缩,且越深越明显。

Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers

  • 将深层残差网络和Transformer的训练转化为无约束特征模型
  • 深度增加时,特征表示趋于神经坍缩,且逼近程度提升
  • 适用于视觉与语言任务,解释了实证现象背后的理论机制

神经坍缩——深度神经网络倒数第二层特征表示中出现的惊人对称性——引发了大量理论研究。然而,现有工作多针对数据无关模型,或仅限于多层感知机。本文填补这两项空白,分析现代架构在数据相关设定下的表现:证明深度正则化变压器和带层归一化的残差网络(ResNets)在交叉熵或均方误差损失下,全局最优解近似呈现神经坍缩,且随着深度增加,逼近程度更优。我们进一步将任意端到端的大深度ResNet或Transformer训练形式化为等价的无约束特征模型,从而解释其在文献中的广泛使用,即便在非数据无关场景下也成立。实验结果在计算机视觉与自然语言处理数据集上验证:随着深度增加,神经坍缩现象愈发显著。

原文摘要 · Abstract (English)

The empirical emergence of neural collapse -- a surprising symmetry in the feature representations of the training data in the penultimate layer of deep neural networks -- has spurred a line of theoretical research aimed at its understanding. However, existing work focuses on data-agnostic models or, when data structure is taken into account, it remains limited to multi-layer perceptrons. Our paper fills both these gaps by analyzing modern architectures in a data-aware regime: we prove that global optima of deep regularized transformers and residual networks (ResNets) with LayerNorm trained with cross entropy or mean squared error loss are approximately collapsed, and the approximation gets tighter as the depth grows. More generally, we formally reduce any end-to-end large-depth ResNet or transformer training into an equivalent unconstrained features model, thus justifying its wide use in the literature even beyond data-agnostic settings. Our theoretical results are supported by experiments on computer vision and language datasets showing that, as the depth grows, neural collapse indeed becomes more prominent.

神经坍缩深度学习变压器残差网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。