arXiv:2502.21269stat.MLcond-mat.dis-nn2025-02NeurIPS被引 29

揭示大模型训练中泛化与过拟合的动态解耦机制

Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks

  • 用动力学平均场理论分析宽层网络训练过程
  • 发现测试误差先降后升,存在特征遗忘阶段
  • 适合研究大模型泛化行为的理论工作者

理解大规模过参数化机器学习模型的归纳偏置和泛化特性,需刻画训练算法的动力学。我们通过动力学平均场理论(一种非平衡统计物理的经典方法),研究大两层神经网络的学习动态。在网络宽度 $m$ 较大、每维样本数 $n/d$ 较大的条件下,训练动态呈现时间尺度分离:(i) 出现与高斯/雷达马克复杂度增长相关的慢时间尺度;(ii) 若初始化复杂度足够小,则产生向低复杂度的归纳偏置;(iii) 特征学习与过拟合阶段实现动态解耦;(iv) 测试误差表现出非单调变化,大时间下出现‘特征遗忘’现象。

原文摘要 · Abstract (English)

Understanding the inductive bias and generalization properties of large overparametrized machine learning models requires to characterize the dynamics of the training algorithm. We study the learning dynamics of large two-layer neural networks via dynamical mean field theory, a well established technique of non-equilibrium statistical physics. We show that, for large network width $m$, and large number of samples per input dimension $n/d$, the training dynamics exhibits a separation of timescales which implies: $(i)$~The emergence of a slow time scale associated with the growth in Gaussian/Rademacher complexity of the network; $(ii)$~Inductive bias towards small complexity if the initialization has small enough complexity; $(iii)$~A dynamical decoupling between feature learning and overfitting regimes; $(iv)$~A non-monotone behavior of the test error, associated `feature unlearning' regime at large times.

神经网络泛化能力动力学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。