arXiv:2509.12467cs.LGcs.NA2025-09

提出非局部神经正切核,让神经网络理论适用于不光滑函数和随机估计器。

Nonlocal Neural Tangent Kernels via Parameter-Space Interactions

  • 用参数空间中的非局部交互替代传统梯度,扩展NTK适用范围
  • 新方法可处理非光滑目标函数与随机优化器,理论覆盖更广
  • 适用于研究不连续或高噪声场景下的模型训练动态

神经正切核(NTK)框架为梯度流下的神经网络训练动态提供了深刻见解。然而,该框架依赖于网络对参数可微的假设,当面对不光滑目标函数或存在不可微行为的参数化模型时,这一假设失效。本文提出非局部神经正切核(NNTK),用参数空间中的非局部交互近似替代局部梯度。非局部梯度存在于比标准梯度更广泛的函数类中,使NTK理论可推广至不光滑函数、随机估计器及更广泛的模型家族。我们探讨了固定核与基于注意力机制的两种非局部算子形式,并通过数值实验验证新方法的有效性。

原文摘要 · Abstract (English)

The Neural Tangent Kernel (NTK) framework has provided deep insights into the training dynamics of neural networks under gradient flow. However, it relies on the assumption that the network is differentiable with respect to its parameters, an assumption that breaks down when considering non-smooth target functions or parameterized models exhibiting non-differentiable behavior. In this work, we propose a Nonlocal Neural Tangent Kernel (NNTK) that replaces the local gradient with a nonlocal interaction-based approximation in parameter space. Nonlocal gradients are known to exist for a wider class of functions than the standard gradient. This allows NTK theory to be extended to nonsmooth functions, stochastic estimators, and broader families of models. We explore both fixed-kernel and attention-based formulations of this nonlocal operator. We illustrate the new formulation with numerical studies.

神经正切核非光滑优化参数空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。