提出深度诱导核,让深层网络理论分析更完整。
Depth-induced NTK: Bridging Over-parameterized Neural Networks and Deep Neural Kernels
- 基于跳跃连接架构构造新核函数,融合深度与无限宽特性。
- 深度无穷时收敛为高斯过程,稳定训练且避免退化。
- 适合研究深度学习理论或核方法的学者参考。
尽管深度学习在众多应用中取得显著成功,但其表征学习的理论理解仍不充分。深度神经核通过将分层特征变换映射到核空间,为过参数化神经网络提供了严谨的解释框架,结合了深度结构的表达力与核方法的可分析性。近期进展,特别是由梯度内积导出的神经正切核(NTK),建立了无限宽神经网络与非参数贝叶斯推断之间的联系。然而,现有NTK范式主要局限于无限宽情形,忽视了网络深度的作用。为此,本文提出一种基于跳跃连接架构的深度诱导神经正切核(Depth-induced NTK),该核在深度趋于无穷时收敛至高斯过程。我们理论分析了该核的训练不变性与谱特性,证明其能稳定核动态并缓解退化问题。实验结果进一步验证了所提方法的有效性。研究显著拓展了神经核理论的边界,深化了对深度学习及缩放规律的理解。
原文摘要 · Abstract (English)
While deep learning has achieved remarkable success across a wide range of applications, its theoretical understanding of representation learning remains limited. Deep neural kernels provide a principled framework to interpret over-parameterized neural networks by mapping hierarchical feature transformations into kernel spaces, thereby combining the expressive power of deep architectures with the analytical tractability of kernel methods. Recent advances, particularly neural tangent kernels (NTKs) derived by gradient inner products, have established connections between infinitely wide neural networks and nonparametric Bayesian inference. However, the existing NTK paradigm has been predominantly confined to the infinite-width regime, while overlooking the representational role of network depth. To address this gap, we propose a depth-induced NTK kernel based on a shortcut-related architecture, which converges to a Gaussian process as the network depth approaches infinity. We theoretically analyze the training invariance and spectrum properties of the proposed kernel, which stabilizes the kernel dynamics and mitigates degeneration. Experimental results further underscore the effectiveness of our proposed method. Our findings significantly extend the existing landscape of the neural kernel theory and provide an in-depth understanding of deep learning and the scaling law.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。