arXiv:2510.15174cs.LG2025-10

揭示神经网络在有限宽度下特征学习的相变机制

A simple mean field model of feature learning

  • 基于平均场理论建模两层网络的特征学习过程
  • 发现有限宽度下存在对目标函数突然对齐的相变现象
  • 提出自增强特征选择机制,解释泛化性能提升

特征学习(FL)指神经网络在训练过程中调整内部表征,仍缺乏深入理解。我们采用统计物理方法,为使用随机梯度朗之万动力学(SGLD)训练的两层非线性网络推导出可解析的、自洽的平均场(MF)理论,描述其贝叶斯后验。在无限宽度下,该理论退化为核岭回归;但在有限宽度下,预测网络会经历对称性破缺相变,突然与目标函数对齐。基础MF理论虽提供有限宽度下特征学习出现的理论洞见,并半定量预测噪声或样本量影响的临界点,但严重低估了相变后的泛化性能提升。我们发现这一偏差源于原理论缺失的关键机制——自增强输入特征选择。将此机制引入MF模型后,能定量匹配SGLD训练网络的学习曲线,并提供特征学习的机制解释。

原文摘要 · Abstract (English)

Feature learning (FL), where neural networks adapt their internal representations during training, remains poorly understood. Using methods from statistical physics, we derive a tractable, self-consistent mean-field (MF) theory for the Bayesian posterior of two-layer non-linear networks trained with stochastic gradient Langevin dynamics (SGLD). At infinite width, this theory reduces to kernel ridge regression, but at finite width it predicts a symmetry breaking phase transition where networks abruptly align with target functions. While the basic MF theory provides theoretical insight into the emergence of FL in the finite-width regime, semi-quantitatively predicting the onset of FL with noise or sample size, it substantially underestimates the improvements in generalisation after the transition. We trace this discrepancy to a key mechanism absent from the plain MF description: \textit{self-reinforcing input feature selection}. Incorporating this mechanism into the MF theory allows us to quantitatively match the learning curves of SGLD-trained networks and provides mechanistic insight into FL.

特征学习平均场深度学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。