arXiv:2410.05626stat.MLcs.LG2024-10NeurIPS被引 7

随机初始化影响神经网络泛化能力,揭示NTK理论的局限性

On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory

  • 分析随机初始化下神经网络的训练动态与NTK回归一致性
  • 证明泛化误差下界为Ω(n^{-3/(d+3)}), 存在维度灾难
  • 指出镜像初始化更优,挑战现有NTK理论解释力

本文探讨随机初始化对神经正切核(NTK)理论的影响,该问题被多数近期研究忽略。当网络宽度趋于无穷时,随机初始化的神经网络收敛于定义在数据域$\mathcal{X}$上的$L^{2}(\mathcal{X})$空间的高斯过程$f^{\mathrm{GP}}$。而多数现有工作采用特殊镜像结构与镜像初始化,使网络初始输出恒为零。本文首先证明:随机初始化下神经网络的梯度流训练动态均匀收敛至对应NTK回归过程。进一步得出:对任意$s < \frac{3}{d+1}$,有$\mathbf{P}(f^{\mathrm{GP}} \in [\mathcal{H}^{\mathrm{NT}}]^{s}) = 1$;对任意$s \geq \frac{3}{d+1}$,则概率为0。由此可得,梯度下降训练的宽网络泛化误差为$Ω(n^{-\frac{3}{d+3}})$,仍受维度灾难制约。结果表明镜像初始化具优势,同时暗示NTK理论可能无法完全解释神经网络的优异性能。

原文摘要 · Abstract (English)

This paper aims to discuss the impact of random initialization of neural networks in the neural tangent kernel (NTK) theory, which is ignored by most recent works in the NTK theory. It is well known that as the network's width tends to infinity, the neural network with random initialization converges to a Gaussian process $f^{\mathrm{GP}}$, which takes values in $L^{2}(\mathcal{X})$, where $\mathcal{X}$ is the domain of the data. In contrast, to adopt the traditional theory of kernel regression, most recent works introduced a special mirrored architecture and a mirrored (random) initialization to ensure the network's output is identically zero at initialization. Therefore, it remains a question whether the conventional setting and mirrored initialization would make wide neural networks exhibit different generalization capabilities. In this paper, we first show that the training dynamics of the gradient flow of neural networks with random initialization converge uniformly to that of the corresponding NTK regression with random initialization $f^{\mathrm{GP}}$. We then show that $\mathbf{P}(f^{\mathrm{GP}} \in [\mathcal{H}^{\mathrm{NT}}]^{s}) = 1$ for any $s < \frac{3}{d+1}$ and $\mathbf{P}(f^{\mathrm{GP}} \in [\mathcal{H}^{\mathrm{NT}}]^{s}) = 0$ for any $s \geq \frac{3}{d+1}$, where $[\mathcal{H}^{\mathrm{NT}}]^{s}$ is the real interpolation space of the RKHS $\mathcal{H}^{\mathrm{NT}}$ associated with the NTK. Consequently, the generalization error of the wide neural network trained by gradient descent is $Ω(n^{-\frac{3}{d+3}})$, and it still suffers from the curse of dimensionality. On one hand, the result highlights the benefits of mirror initialization. On the other hand, it implies that NTK theory may not fully explain the superior performance of neural networks.

NTK理论泛化误差随机初始化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。