arXiv:2505.06597cs.LGcond-mat.dis-nn2025-05

L2正则化引发深度模型精度相变,解释了‘领悟’现象的本质。

Phase Transitions between Accuracy Regimes in L2 regularized Deep Neural Networks

  • 用误差曲面的标量曲率解释正则化导致的相变机制
  • 预测数据复杂度增加时新相变点存在,且出现滞后效应
  • 为理解模型局部极小与学习行为提供新视角,适合研究模型优化者

增加深度神经网络(DNN)的L2正则化会引发一次一阶相变,进入欠参数化相——即所谓的“学习启动点”。我们通过误差景观的标量(Ricci)曲率解释这一相变。随着数据复杂度提升,我们预测了新的相变点,并根据相变理论预测存在滞后效应,数值实验验证了这两项预测。我们的结果为近期发现的‘领悟’(grokking)现象提供了自然解释:模型陷入误差表面的局部极小,对应较低准确率相。本工作为深入探究DNN内在结构(超越L2正则化)开辟了新路径。

原文摘要 · Abstract (English)

Increasing the L2 regularization of Deep Neural Networks (DNNs) causes a first-order phase transition into the under-parametrized phase -- the so-called onset-of learning. We explain this transition via the scalar (Ricci) curvature of the error landscape. We predict new transition points as the data complexity is increased and, in accordance with the theory of phase transitions, the existence of hysteresis effects. We confirm both predictions numerically. Our results provide a natural explanation of the recently discovered phenomenon of '\emph{grokking}' as DNN models getting stuck in a local minimum of the error surface, corresponding to a lower accuracy phase. Our work paves the way for new probing methods of the intrinsic structure of DNNs in and beyond the L2 context.

深度学习正则化相变优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。