arXiv:2607.04135cond-mat.dis-nncond-mat.stat-mech2026-07被引 1

神经网络容量越大越能泛化,这源于训练过程中的相变与对称性破缺。

Broken Ergodicity and the Violation of the Fluctuation-Dissipation Theorem Lead to Generalization Beyond Overfitting in Machine Learning

  • 用动态平均场理论揭示训练过程存在相变,解释双下降现象
  • 发现泛化能力在临界点附近突增,与超导伦敦模型形式相同
  • 适合研究深度学习泛化机制的理论学者参考

现代神经网络的泛化能力随模型容量增加而提升,即使参数量超过训练数据点数。这一现象尤为反常,因为当参数量接近临界值时,泛化误差会发散。我们通过动态平均场理论证明,这种所谓的‘双下降’行为是描述训练过程的随机场论中相变的结果。我们计算了双下降相变的临界指数和标度函数,并发现其特征为涨落-耗散定理的破坏,源于对称性破缺。对应的响应函数与超导转变的简单伦敦模型形式一致,波函数刚度对应神经网络的准确泛化能力。

原文摘要 · Abstract (English)

The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points. This phenomenon is all the more surprising given that generalization error diverges when the number of model parameters approaches a critical value from below. Here we use dynamical mean field theory to show that this so-called "double descent" behavior is the outcome of a phase transition in the stochastic field theory describing the training process. We calculate the critical exponents and scaling function of the double descent phase transition, and show that it is marked by a breakdown of the fluctuation-dissipation theorem associated with broken ergodicity. The corresponding response function has the same functional form as the simple London model of the superconducting transition, with the rigidity of the wave function corresponding to the neural network's ability to generalize accurately.

泛化能力相变神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。