arXiv:2502.07998cs.LGcond-mat.dis-nn2025-02ICML被引 19

神经网络在无限宽下也能用自适应核函数描述,提升模型性能。

Adaptive kernel predictors from feature-learning infinite limits of neural networks

  • 通过贝叶斯网络大宽极限推导出任务相关的自适应核函数。
  • 在梯度流训练中,自适应核比懒惰模式核更优,测试误差更低。
  • 适用于研究深度学习泛化机制或设计新核方法的研究者。

以往重要工作表明,在懒惰训练模式下,神经网络的无限宽度极限可由核机器描述。本文发现,在两种不同设置下的丰富特征学习无限宽度模式中,神经网络同样可由数据依赖的核机器描述。我们给出了两种核预测器的显式表达式,并提供数值计算方法。第一种基于特征学习贝叶斯网络的大宽极限,揭示特征学习如何使层核和预激活分布适应任务需求;其鞍点方程构成一个极小极大优化问题,定义了核预测器。第二种基于随机初始化网络在权重衰减下的梯度流训练,利用动力学平均场理论(DMFT)分析无限宽度极限,其固定点方程定义了任务自适应的内部表示与核预测器。我们将所得自适应核与懒惰模式核对比,结果表明:在基准数据集上,自适应核达到更低的测试损失。

原文摘要 · Abstract (English)

Previous influential work showed that infinite width limits of neural networks in the lazy training regime are described by kernel machines. Here, we show that neural networks trained in the rich, feature learning infinite-width regime in two different settings are also described by kernel machines, but with data-dependent kernels. For both cases, we provide explicit expressions for the kernel predictors and prescriptions to numerically calculate them. To derive the first predictor, we study the large-width limit of feature-learning Bayesian networks, showing how feature learning leads to task-relevant adaptation of layer kernels and preactivation densities. The saddle point equations governing this limit result in a min-max optimization problem that defines the kernel predictor. To derive the second predictor, we study gradient flow training of randomly initialized networks trained with weight decay in the infinite-width limit using dynamical mean field theory (DMFT). The fixed point equations of the arising DMFT defines the task-adapted internal representations and the kernel predictor. We compare our kernel predictors to kernels derived from lazy regime and demonstrate that our adaptive kernels achieve lower test loss on benchmark datasets.

神经网络核方法特征学习无限宽度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。