arXiv:2603.13872cs.LGmath.DS2026-03被引 1

揭示神经网络泛化机制,证明训练过程如内核机器般记忆特征。

On Interpolation Formulas Describing Neural Network Generalization

  • 提出随机梯度内核,将优化器权重纳入泛化建模。
  • 发现测试点与学习特征记忆的对齐决定泛化性能。
  • 统一解释扩散模型与GAN的修正机制,适合研究泛化理论者。

2020年Domingos提出适用于“所有梯度下降训练模型”的插值公式,推断此类模型近似为核机器。本文将该公式扩展至随机训练,引入通过连续时间扩散近似构造的随机梯度核。我们证明了随机版Domingos定理,表明期望网络输出可表示为与优化器相关的核机器形式,训练样本通过损失相关权重和梯度对齐贡献。我们进一步将泛化误差关联到随机梯度核诱导的积分算子零空间。相同的路径核视角统一解释了扩散模型与GAN:扩散实现分阶段、噪声局部化的修正,而GAN则受判别器几何引导的分布修正。我们可视化优化过程中隐式核的演化,并通过一系列数值实验量化了分布外行为。结果支持特征空间记忆学习观:训练在动态切线特征几何中存储数据相关信息,测试时预测源于核加权的特征检索与聚合,泛化由测试点与学习特征记忆间的对齐决定。

原文摘要 · Abstract (English)

In 2020 Domingos introduced an interpolation formula valid for "every model trained by gradient descent". He concluded that such models behave approximately as kernel machines. In this work, we extend the Domingos formula to stochastic training. We introduce a stochastic gradient kernel that extends the deterministic version via a continuous-time diffusion approximation. We prove stochastic Domingos theorems and show that the expected network output admits a kernel-machine representation with optimizer-specific weighting. It reveals that training samples contribute through loss-dependent weights and gradient alignment along the training trajectory. We then link the generalization error to the null space of the integral operator induced by the stochastic gradient kernel. The same path-kernel viewpoint provides a unified interpretation of diffusion models and GANs: diffusion induces stage-wise, noise-localized corrections, whereas GANs induce distribution-guided corrections shaped by discriminator geometry. We visualize the evolution of implicit kernels during optimization and quantify out-of-distribution behaviors through a series of numerical experiments. Our results support a feature-space memory view of learning: training stores data-dependent information in an evolving tangent feature geometry, and predictions at test time arise from kernel-weighted retrieval and aggregation of these stored features, with generalization governed by alignment between test points and the learned feature memory.

泛化分析核方法优化器特征记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。