arXiv:2410.23169cs.LG2024-10NeurIPS被引 14

证明深度网络中神经坍缩现象仍普遍存在,即使在交叉熵损失下也难被低秩偏置消除。

The Persistence of Neural Collapse Despite Low-Rank Bias

  • 扩展理论至交叉熵损失下的深层无约束特征模型,揭示高秩结构非最优
  • 证明全局最小值处非零奇异值数量存在固定上限,随网络加深不增
  • 解释神经坍缩频现于实证:其在损失曲面中更易被找到,具统计优势

神经坍缩(NC)及其多层变体深度神经坍缩(DNC)描述了训练后深层网络特征与权重的结构性几何。近期由Sukenik等人基于深层无约束特征模型(UFM)的理论工作表明,在均方误差(MSE)损失下,DNC并非最优。他们推测这是由L2正则化引发的低秩偏置所致。本文将该结论扩展至使用交叉熵损失训练的深层UFM,证明高秩结构(包括DNC)通常非最优。我们刻画了相关低秩偏置,证明当网络深度增加时,全局最小值处非零奇异值数量存在固定上界。进一步分析损失曲面,发现相较于其他临界配置,DNC在优化景观中更为普遍,这解释了其在实验中的频繁出现。结果通过深层UFM和深层神经网络的实验得到验证。

原文摘要 · Abstract (English)

Neural collapse (NC) and its multi-layer variant, deep neural collapse (DNC), describe a structured geometry that occurs in the features and weights of trained deep networks. Recent theoretical work by Sukenik et al. using a deep unconstrained feature model (UFM) suggests that DNC is suboptimal under mean squared error (MSE) loss. They heuristically argue that this is due to low-rank bias induced by L2 regularization. In this work, we extend this result to deep UFMs trained with cross-entropy loss, showing that high-rank structures, including DNC, are not generally optimal. We characterize the associated low-rank bias, proving a fixed bound on the number of non-negligible singular values at global minima as network depth increases. We further analyze the loss surface, demonstrating that DNC is more prevalent in the landscape than other critical configurations, which we argue explains its frequent empirical appearance. Our results are validated through experiments in deep UFMs and deep neural networks.

神经坍缩深度学习优化理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。