arXiv:2510.21078cs.LGmath.OC2025-10NeurIPS被引 2

证明了浅层ReLU网络在梯度流下可实现神经坍缩,揭示数据结构与激活函数的作用。

Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data

  • 在正交可分数据上研究两层ReLU网络的梯度流优化
  • 证明训练过程必然导致神经坍缩现象出现
  • 揭示了训练动态隐式偏差对神经坍缩的促进作用

深度网络成功背后的一个谜团是其学习表征所展现出的惊人判别力,这由令人瞩目的神经坍缩(Neural Collapse, NC)现象体现:训练后网络最后一层的特征会呈现简洁结构。此前理论研究多将最后一层特征视为无约束自由变量,通过分析矩阵分解类问题的优化景观,证明全局极小点具有NC特性。本文首次证明,在对正交可分数据进行分类时,两层ReLU网络在梯度流下的优化过程可严格实现神经坍缩。这一结果推进了已有成果:第一,放宽了特征无约束的假设,揭示了数据结构与非线性激活函数对NC特征的影响;第二,阐明了训练动态的隐式偏差在促成NC涌现中的关键作用。

原文摘要 · Abstract (English)

Among many mysteries behind the success of deep networks lies the exceptional discriminative power of their learned representations as manifested by the intriguing Neural Collapse (NC) phenomenon, where simple feature structures emerge at the last layer of a trained neural network. Prior works on the theoretical understandings of NC have focused on analyzing the optimization landscape of matrix-factorization-like problems by considering the last-layer features as unconstrained free optimization variables and showing that their global minima exhibit NC. In this paper, we show that gradient flow on a two-layer ReLU network for classifying orthogonally separable data provably exhibits NC, thereby advancing prior results in two ways: First, we relax the assumption of unconstrained features, showing the effect of data structure and nonlinear activations on NC characterizations. Second, we reveal the role of the implicit bias of the training dynamics in facilitating the emergence of NC.

神经坍缩深度学习理论梯度流ReLU网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。