arXiv:2605.17718stat.MLcs.LG2026-05

揭示神经网络特征学习如何重塑函数空间,使关键方向被优先增强。

How does feature learning reshape the function space?

  • 通过梯度下降分析两层网络的特征分布演化,发现其呈现目标相关的尖峰高斯结构。
  • 特征学习使函数空间谱结构改变,显著放大与目标对齐的特征方向。
  • 适合研究深度学习优化机制或函数空间理论的学者参考。

特征学习被视为神经网络区别于固定核方法的核心机制,但其对函数空间的影响仍不清晰。本文精确刻画了两层神经网络在梯度下降过程中特征所张成的函数空间演变。在高维比例条件下,经过大量梯度步后,更新后的特征分布可被目标相关的尖峰高斯协方差良好近似,从而诱导出一种数据自适应核,重塑函数空间并改变其谱结构。我们的分析表明,特征学习可被理解为参数空间或输入空间中的分布变换,等价于引入目标依赖的核。具体而言,它选择性放大与目标方向对齐的特征值,并混合主导特征函数,将顶层径向模式与目标对齐的二次调和函数耦合。总体上,本研究提供了早期特征学习在函数空间层面的精确视角:梯度下降并非仅缩放固定核,而是诱导出一种数据自适应的形变,优先增强数据中信号对齐的方向。

原文摘要 · Abstract (English)

Feature learning is widely regarded as the key mechanism distinguishing neural networks from fixed-kernel methods, yet its impact on the induced function space remains poorly understood. In this work, we precisely characterize how the function space spanned by the features of a two-layer neural network evolves during gradient descent training. We prove that, in the high-dimensional proportional regime, after a large gradient step the post-update feature distribution is well approximated by a target-dependent spiked Gaussian covariance. This induces a data-adaptive kernel that reshapes the function space and modifies its spectral structure. Our analysis reveals that feature learning can be interpreted as a distributional transformation in either parameter space or input space, equivalently as the introduction of a target-dependent kernel. In particular, it selectively amplifies eigenvalues aligned with the target direction and mixes leading eigenfunctions, coupling the top radial mode with a target-aligned quadratic harmonic. Overall, our results provide a precise function-space perspective on early-stage feature learning: rather than just rescaling a fixed kernel, gradient descent induces a data-adaptive deformation that preferentially enhances directions aligned with the signal in the data.

函数空间特征学习神经网络优化谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。