揭示标签与核矩阵的对齐机制,提升神经网络收敛预测精度。
Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime

- 发现标签与NTK特征向量的对齐现象,影响训练速度。
- 新边界比传统结果快10倍以上,更贴近实际训练过程。
- 适用于理解大模型优化,适合研究者与工程师参考。
神经正切核(NTK)框架通过近似线性动力学解释过参数化神经网络的优化,提供指数收敛保证。但现有结果通常过于保守,未能反映实际训练中的快速收敛,因其依赖于最小的NTK特征值,而该值在实践中通常极小。本文通过刻画数据标签与NTK特征谱的相互作用,提出更紧致的收敛界。我们识别出两个关键现象:标签-NTK对齐与残差-NTK对齐,表明标签和残差在NTK特征向量上的投影大小与对应特征值成正比。在弱数据假设下提供理论支持与实证证据。利用这些对齐性质,我们推导出依赖于全谱的改进收敛界,与实际训练动态高度吻合,显著优于经典最坏情况结果。进一步获得改进的泛化界。在MLP与CNN上,多个数据集的实验验证了理论的有效性。
原文摘要 · Abstract (English)
The Neural Tangent Kernel (NTK) framework explains optimization in over-parameterized neural networks via approximately linearized dynamics, yielding exponential convergence guarantees. However, existing results are often overly pessimistic and do not match the fast training in practice, as they depend on the smallest NTK eigenvalue, which is typically extremely small in practice. In this work, we develop sharper convergence guarantees by characterizing the interaction between data labels and the NTK eigen-spectrum. We identify two key phenomena, Label-NTK alignment and Residual-NTK alignment, showing that projections of labels and residuals onto NTK eigenvectors scale with the corresponding eigenvalues. We provide empirical evidence and theoretical justification under mild data assumptions. Exploiting these alignment properties, we derive a refined convergence bound that depends on the full spectrum and closely matches practical training dynamics, significantly improving over classical worst-case results. We further obtain improved generalization bounds. Experiments on MLPs and CNNs across multiple datasets validate our theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。