arXiv:2409.01832stat.MLcs.LG2024-09被引 9

揭示浅层神经网络中特征坍缩的形成条件及其与数据质量的关系

Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data

  • 分析浅层ReLU网络在不同宽度、深度下的特征坍缩现象
  • 发现特征坍缩取决于数据维度、样本量和信噪比,而非网络宽度
  • 揭示低信噪比下即使发生坍缩也无法保证良好泛化

神经坍缩(Neural Collapse, NC)是深度神经网络训练末期出现的现象:同类样本特征趋同于其均值,且类均值构成等角紧框架(ETF)。以往研究多基于无约束特征模型(UFM),难以刻画网络结构与数据特性对NC的影响。本文聚焦浅层ReLU网络,系统分析宽度、深度、数据维度及统计特性对NC出现条件的影响。对于两层网络,证明全局极小解呈现NC仅依赖于数据维度、样本数和信噪比(SNR),与网络宽度无关;三层数时,只要第一层足够宽,即出现NC。此外,发现泛化性能高度依赖数据中SNR:即使存在NC,若数据信噪比过低,泛化仍差。本工作首次在非受限条件下完整刻画了浅层非线性网络中NC的产生机制及其与数据和架构的关系。

原文摘要 · Abstract (English)

Neural collapse (NC) is a phenomenon that emerges at the terminal phase of the training (TPT) of deep neural networks (DNNs). The features of the data in the same class collapse to their respective sample means and the sample means exhibit a simplex equiangular tight frame (ETF). In the past few years, there has been a surge of works that focus on explaining why the NC occurs and how it affects generalization. Since the DNNs are notoriously difficult to analyze, most works mainly focus on the unconstrained feature model (UFM). While the UFM explains the NC to some extent, it fails to provide a complete picture of how the network architecture and the dataset affect NC. In this work, we focus on shallow ReLU neural networks and try to understand how the width, depth, data dimension, and statistical property of the training dataset influence the neural collapse. We provide a complete characterization of when the NC occurs for two or three-layer neural networks. For two-layer ReLU neural networks, a sufficient condition on when the global minimizer of the regularized empirical risk function exhibits the NC configuration depends on the data dimension, sample size, and the signal-to-noise ratio in the data instead of the network width. For three-layer neural networks, we show that the NC occurs as long as the first layer is sufficiently wide. Regarding the connection between NC and generalization, we show the generalization heavily depends on the SNR (signal-to-noise ratio) in the data: even if the NC occurs, the generalization can still be bad provided that the SNR in the data is too low. Our results significantly extend the state-of-the-art theoretical analysis of the N C under the UFM by characterizing the emergence of the N C under shallow nonlinear networks and showing how it depends on data properties and network architecture.

神经坍缩浅层网络信噪比泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。