arXiv:2502.01556cs.LGstat.ML2025-02被引 6

将宽神经网络的训练映射为带噪声观测的高斯过程,提升实际应用性。

A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks

  • 引入正则化项模拟观测噪声,修正原模型对噪声数据的误设
  • 提出移位网络结构,支持任意先验均值,无需集成或核反演
  • 实验验证方法在多种数据集和架构上有效,推动NTK-GP实用化

在宽神经网络中进行梯度下降等价于计算具有神经正切核(NTK-GP)的高斯过程后验均值,但该框架存在两方面限制:(i) NTK-GP假设目标无噪声,导致在有噪声数据上出现模型误设;(ii) 等价性不适用于任意先验均值,而后者对模型正确设定至关重要。为解决(i),我们在训练目标中引入正则化项,证明其对应于在NTK-GP中加入观测噪声。为解决(ii),我们提出一种“移位网络”结构,可实现任意先验均值,并通过单个网络的梯度下降即可获得后验均值,无需集成或核反演。我们在多个数据集和网络架构上进行了实验验证,结果表明该方法消除了应用中使用NTK-GP等价性的关键障碍。

原文摘要 · Abstract (English)

Performing gradient descent in a wide neural network is equivalent to computing the posterior mean of a Gaussian Process with the Neural Tangent Kernel (NTK-GP), for a specific prior mean and with zero observation noise. However, existing formulations have two limitations: (i) the NTK-GP assumes noiseless targets, leading to misspecification on noisy data; (ii) the equivalence does not extend to arbitrary prior means, which are essential for well-specified models. To address (i), we introduce a regularizer into the training objective, showing its correspondence to incorporating observation noise in the NTK-GP. To address (ii), we propose a \textit{shifted network} that enables arbitrary prior means and allows obtaining the posterior mean with gradient descent on a single network, without ensembling or kernel inversion. We validate our results with experiments across datasets and architectures, showing that this approach removes key obstacles to the practical use of NTK-GP equivalence in applied Gaussian process modeling.

高斯过程神经网络核方法贝叶斯深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。