突破性扩展神经正切核理论至分类任务,揭示其在特定条件下仍保持稳定。
The Neural Tangent Kernel for Classification
- 通过参数正则化或非退化标签,确保分类任务中神经正切核恒定不变。
- 在上述条件下,训练可由线性模型精确近似,解可显式表示为神经正切核函数。
- 关联随机初始化下的预测分布与贝叶斯方法,提供模型不确定性新视角。
在宽神经网络中,神经正切核(NTK)在训练过程中近似恒定,为研究训练动态、泛化能力及与核方法的联系提供了有力理论工具。然而,该理论主要局限于回归损失。此前认为,使用分类损失或涉及非线性输出变换的损失会破坏此性质,导致对数输出发散,线性化失效。本文通过识别条件,将NTK理论扩展至分类任务:参数空间正则化可保证交叉熵损失下训练期间NTK恒定;若无正则化,则当标签非退化(即所有类别概率严格为正)时,该性质仍成立。在此条件下,训练可被线性模型良好近似,解可显式表征为神经正切核形式。我们进一步分析了由随机初始化诱导的训练预测器分布,并将其与贝叶斯方法中的模型不确定性相联系。
原文摘要 · Abstract (English)
In wide neural networks, the Neural Tangent Kernel (NTK) remains approximately constant during training, providing a powerful theoretical tool for studying training dynamics, generalization, and connections to kernel methods. However, this theory is largely restricted to regression losses. It was previously thought that training on a classification loss, or more generally losses involving nonlinear output transformations, breaks this property, leading to divergent logits and a breakdown of the linearization. In this paper, we extend NTK theory to classification by identifying conditions under which wide neural networks remain in the lazy training regime. We show that parameter-space regularization ensures a constant NTK during training for cross-entropy loss, while in the absence of regularization the regime is recovered when targets are non-degenerate, i.e. when all classes have strictly positive probability. Under these conditions, training is well-approximated by the linearized model, yielding an explicit characterization of the solution in terms of the NTK. We further analyze the distribution of trained predictors induced by random initialization and relate this notion of model uncertainty to Bayesian methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。