深度宽度同比无限增长时,线性神经网络收敛到非高斯混合分布。
Proportional infinite-width infinite-depth limit for deep linear neural networks
- 深度与宽度按比例同步无限增长,突破传统固定深度极限。
- 得到非高斯混合分布,能保留输出间的相关性特征。
- 适用于需要建模复杂依赖关系的深层网络理论分析。
我们研究了在大网络背景下随机参数线性神经网络的分布特性,其中层数与每层神经元数以相同比例发散。以往研究显示,在无限宽极限下(层数固定,每层神经元数趋于无穷),神经网络收敛于高斯过程,即神经网络高斯过程(Neural Network Gaussian Process)。然而,该高斯极限因无法学习相关特征、无法生成反映真实标签的相关输出而损失描述能力。为克服此局限,我们探索了深度与宽度同步发散但保持恒定比例的联合比例极限,得到一个保留输出相关性的非高斯分布。我们的贡献在于严格刻画了线性激活函数下该极限分布为非平凡的高斯混合分布。
原文摘要 · Abstract (English)
We study the distributional properties of linear neural networks with random parameters in the context of large networks, where the number of layers diverges in proportion to the number of neurons per layer. Prior works have shown that in the infinite-width regime, where the number of neurons per layer grows to infinity while the depth remains fixed, neural networks converge to a Gaussian process, known as the Neural Network Gaussian Process. However, this Gaussian limit sacrifices descriptive power, as it lacks the ability to learn dependent features and produce output correlations that reflect observed labels. Motivated by these limitations, we explore the joint proportional limit in which both depth and width diverge but maintain a constant ratio, yielding a non-Gaussian distribution that retains correlations between outputs. Our contribution extends previous works by rigorously characterizing, for linear activation functions, the limiting distribution as a nontrivial mixture of Gaussians.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。