arXiv:2505.11347cs.LG2025-05被引 1

让NTK直接优化泛化误差,比传统训练更有效。

Training NTK to Generalize with KARE

  • 用KARE直接优化NTK的泛化误差,不依赖原始网络训练。
  • 实验证明新方法在多个数据集上超越原DNN和后导NTK。
  • 适合关注泛化性能与特征学习机制的研究者。

深度神经网络(DNN)训练过程中,其对应的依赖数据的神经正切核(NTK;Jacot et al. (2018))性能常与或超过完整网络。这表明梯度下降隐式地通过优化NTK实现核学习。本文提出显式优化NTK:不再最小化经验风险,而是利用最近提出的核对齐风险估计器(KARE;Jacot et al. (2020))来最小化NTK的泛化误差。仿真与真实数据实验显示,使用KARE训练的NTK始终匹配或显著优于原始DNN及由DNN诱导出的后导NTK(after-kernel)。结果表明,在某些场景下,显式训练的核可超越传统的端到端DNN优化,挑战了DNN的主导地位。我们认为,显式训练NTK是一种过参数化特征学习形式。

原文摘要 · Abstract (English)

The performance of the data-dependent neural tangent kernel (NTK; Jacot et al. (2018)) associated with a trained deep neural network (DNN) often matches or exceeds that of the full network. This implies that DNN training via gradient descent implicitly performs kernel learning by optimizing the NTK. In this paper, we propose instead to optimize the NTK explicitly. Rather than minimizing empirical risk, we train the NTK to minimize its generalization error using the recently developed Kernel Alignment Risk Estimator (KARE; Jacot et al. (2020)). Our simulations and real data experiments show that NTKs trained with KARE consistently match or significantly outperform the original DNN and the DNN- induced NTK (the after-kernel). These results suggest that explicitly trained kernels can outperform traditional end-to-end DNN optimization in certain settings, challenging the conventional dominance of DNNs. We argue that explicit training of NTK is a form of over-parametrized feature learning.

NTK泛化误差核学习KARE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。