无穷宽深层网络在μP参数化下可实现全局收敛与丰富特征学习
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
- 采用μP参数化与随机梯度下降,使特征显著偏离初始值
- 训练后特征线性独立,所有收敛点均为全局最优
- 理论结合实验验证,适用于深度表征学习研究者
尽管深度神经网络具备强大的表征学习能力,但对其如何同时实现有意义的特征学习和全局收敛的理论理解仍不充分。现有方法如神经正切核(NTK)受限于特征始终接近初始化状态。本文基于张量程序(TP)框架,研究在最大更新参数化(μP)下,使用随机梯度下降(SGD)训练的无穷宽$L$层神经网络的动态行为。在激活函数满足温和条件时,我们证明:网络能学习到显著偏离初始值、线性独立的特征,该丰富特征空间捕获了数据的关键信息,并确保训练过程的任意收敛点均为全局最小值。分析结合了层间特征交互与高斯随机变量性质,为深度表征学习提供了新视角。实验在真实数据集上验证了理论结果。
原文摘要 · Abstract (English)
Despite deep neural networks' powerful representation learning capabilities, theoretical understanding of how networks can simultaneously achieve meaningful feature learning and global convergence remains elusive. Existing approaches like the neural tangent kernel (NTK) are limited because features stay close to their initialization in this parametrization, leaving open questions about feature properties during substantial evolution. In this paper, we investigate the training dynamics of infinitely wide, $L$-layer neural networks using the tensor program (TP) framework. Specifically, we show that, when trained with stochastic gradient descent (SGD) under the Maximal Update parametrization ($μ$P) and mild conditions on the activation function, SGD enables these networks to learn linearly independent features that substantially deviate from their initial values. This rich feature space captures relevant data information and ensures that any convergent point of the training process is a global minimum. Our analysis leverages both the interactions among features across layers and the properties of Gaussian random variables, providing new insights into deep representation learning. We further validate our theoretical findings through experiments on real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。