分析了神经网络在过参数化下的牛顿法收敛性,揭示其可快速拟合高频数据。
Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit
- 引入牛顿神经正切核(NNTK),描述过参数化下训练动态的极限行为
- 证明无限宽时模型指数级收敛至零损失,且对高频数据无偏差
- 提出正则化参数调节公式,确保训练中海森矩阵正定,适合研究优化机制
本文针对过参数化极限下的正则化牛顿法训练神经网络,建立了收敛性分析。当隐藏单元数量趋于无穷时,神经网络训练动态以概率收敛到一个确定性极限方程,该方程涉及“牛顿神经正切核”(NNTK)。文中给出了刻画此收敛速度的显式速率,并在无限宽度极限下证明神经网络以指数速度收敛至目标数据(即零损失全局最小值)。该收敛在频率谱上一致,克服了梯度下降固有的频谱偏见:梯度下降的NTK特征值趋于零,导致高频数据收敛缓慢;而适当的正则化下,NNTK特征值有统一下界,使牛顿法能更快处理高频成分。分析面临两大数学挑战:牛顿法隐式更新及海森矩阵可能不定,以及随网络宽度增加线性系统维数趋于无穷。这些使得在过参数化极限下推导训练动态并证明有限宽度动态收敛变得复杂。本文还提出了正则化参数的缩放公式,表明其可随隐藏单元数增大以适当速率趋近于零。我们证明:当隐藏单元数足够大时,正则化海森矩阵在整个训练过程中保持正定,且单个网络参数的牛顿更新趋于零,说明模型行为可线性化于初始化点。
原文摘要 · Abstract (English)
A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. As the number of hidden units tends to infinity, the NN training dynamics converge in probability to the solution of a deterministic limit equation involving a ``Newton neural tangent kernel'' (NNTK). Explicit rates characterizing this convergence are provided and, in the infinite-width limit, we prove that the NN converges exponentially fast to the target data (i.e., a global minimizer with zero loss). We show that this convergence is uniform across the frequency spectrum, addressing the spectral bias inherent in gradient descent. The eigenvalues of the NTK for gradient descent accumulate at zero, leading to slow convergence for target data with high-frequency components. In contrast, the NNTK has uniformly lower bounded eigenvalues if the regularization parameter is selected appropriately, allowing Newton's method to converge more quickly for data with high-frequency components. Mathematical challenges that need to be addressed in our analysis include the implicit parameter update of the Newton method with a potentially indefinite Hessian matrix and the fact that the dimension of this linear system of equations tends to infinity as the NN width grows. This complicates deriving the training dynamics in the overparameterized limit as well as proving the convergence of the finite-width dynamics thereto. The analysis identifies a scaling formula for selecting the regularization parameter, which we show can vanish at a suitable rate as the number of hidden units becomes larger. We prove that, for sufficiently large numbers of hidden units, the regularized Hessian remains positive definite during training and the Newton updates for individual NN parameters converge to zero, showing that the model behaves as a linearization around the initialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。