arXiv:2501.02364cs.LGcs.CV2025-01被引 3

浅层非线性网络可将低维数据映射为线性可分,宽度仅需随内在维度多项式增长。

Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension

  • 用随机权重和二次激活函数,单层网络实现数据线性可分
  • 网络宽度只需与数据内在维度的多项式相关,不依赖高维空间
  • 理论结合实验验证,为模型可解释性提供新视角

深度神经网络在各类分类任务中表现卓越。近期实证研究发现,深度网络学习到的特征具有类间线性可分性,但缺乏严谨理论支撑,尤其在简单设定下。本文聚焦浅层非线性网络的线性分离能力,基于图像数据的低内在维度特性,将输入建模为低维子空间并集(UoS)。理论证明:当使用随机权重和二次激活函数时,单层网络可在高概率下将此类数据映射为线性可分集合。关键结果表明,该变换在网络宽度仅随数据内在维度多项式增长时即可实现,而非依赖环境维度。实验结果支持理论结论,并显示此类线性可分性质在超出分析范围的实际场景中依然成立。本工作弥合了经验观察与理论理解之间的差距,深化了对非线性网络分离能力、模型可解释性与泛化性能的认识。

原文摘要 · Abstract (English)

Deep neural networks have attained remarkable success across diverse classification tasks. Recent empirical studies have shown that deep networks learn features that are linearly separable across classes. However, these findings often lack rigorous justifications, even under relatively simple settings. In this work, we address this gap by examining the linear separation capabilities of shallow nonlinear networks. Specifically, inspired by the low intrinsic dimensionality of image data, we model inputs as a union of low-dimensional subspaces (UoS) and demonstrate that a single nonlinear layer can transform such data into linearly separable sets. Theoretically, we show that this transformation occurs with high probability when using random weights and quadratic activations. Notably, we prove this can be achieved when the network width scales polynomially with the intrinsic dimension of the data rather than the ambient dimension. Experimental results corroborate these theoretical findings and demonstrate that similar linear separation properties hold in practical scenarios beyond our analytical scope. This work bridges the gap between empirical observations and theoretical understanding of the separation capacity of nonlinear networks, offering deeper insights into model interpretability and generalization.

神经网络线性可分浅层网络内在维度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。