揭示多项式神经网络的可识别性与奇异点,解释了 MLP 的稀疏偏好。
Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks
- 用代数几何分析神经网络函数空间,证明多数函数对应有限参数解
- MLP 神经流形维度可计算,CNN 参数映射几乎一一对应
- 发现奇异点源于稀疏子网,为 MLP 的稀疏性提供几何解释
我们研究由神经网络参数化的函数空间(即神经流形)。针对深层多层感知机(MLP)和卷积神经网络(CNN),激活函数为充分通用的多项式。首先解决可识别性问题:对 MLP 中几乎所有函数,仅存在有限多个参数配置能生成该函数;对 CNN,参数化是典型的单射。由此可计算神经流形的维度。其次,完整刻画了 CNN 的奇异点,部分刻画了 MLP 的奇异点。两类奇异点均源于稀疏子网络。对于 MLP,我们证明这些奇异点常对应均方误差损失的临界点,而此性质不适用于 CNN。这为 MLP 的稀疏偏好提供了几何解释。所有结论基于代数几何工具。
原文摘要 · Abstract (English)
We study function spaces parametrized by neural networks, referred to as neuromanifolds. Specifically, we focus on deep Multi-Layer Perceptrons (MLPs) and Convolutional Neural Networks (CNNs) with an activation function that is a sufficiently generic polynomial. First, we address the identifiability problem, showing that, for almost all functions in the neuromanifold of an MLP, there exist only finitely many parameter choices yielding that function. For CNNs, the parametrization is generically one-to-one. As a consequence, we compute the dimension of the neuromanifold. Second, we describe singular points of neuromanifolds. We characterize singularities completely for CNNs, and partially for MLPs. In both cases, they arise from sparse subnetworks. For MLPs, we prove that these singularities often correspond to critical points of the mean-squared error loss, which does not hold for CNNs. This provides a geometric explanation of the sparsity bias of MLPs. All of our results leverage tools from algebraic geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。