arXiv:2412.09810cs.LG2024-12被引 24

发现神经网络在长期训练后会突然从记忆转向泛化,且复杂度有相变现象。

The Complexity Dynamics of Grokking

  • 用率失真理论和可计算复杂度定义网络复杂度,实现有原理的压缩。
  • 正则化网络在训练后期复杂度骤降,体现对简单模式的发现与泛化。
  • 提出基于谱熵的正则化方法,可引导模型向低复杂度表示学习。

我们通过研究神经网络中的grokking现象,揭示了其存在复杂度相变:网络在过度拟合训练数据后,仍会在长期训练中突然从记忆模式转向泛化。为此,我们基于率失真理论和柯尔莫戈洛夫复杂度,提出一个理论框架来衡量网络复杂度,可视为面向网络的有原则的有损压缩。我们发现,适当正则化的网络表现出清晰的相变:复杂度在记忆阶段上升,随后下降,反映出对更简单底层规律的发现;而未正则化网络则陷入高复杂度的记忆状态。我们建立了该复杂度度量与泛化界之间的显式联系,为有损压缩与泛化之间的关联提供了理论基础。所提框架实现30-40倍于朴素方法的压缩比,能精确追踪复杂度动态变化。最后,我们引入一种基于谱熵的正则化方法,通过惩罚内在维度来鼓励网络学习低复杂度表示。

原文摘要 · Abstract (English)

We demonstrate the existence of a complexity phase transition in neural networks by studying the grokking phenomenon, where networks suddenly transition from memorization to generalization long after overfitting their training data. To characterize this phase transition, we introduce a theoretical framework for measuring complexity based on rate-distortion theory and Kolmogorov complexity, which can be understood as principled lossy compression for networks. We find that properly regularized networks exhibit a sharp phase transition: complexity rises during memorization, then falls as the network discovers a simpler underlying pattern that generalizes. In contrast, unregularized networks remain trapped in a high-complexity memorization phase. We establish an explicit connection between our complexity measure and generalization bounds, providing a theoretical foundation for the link between lossy compression and generalization. Our framework achieves compression ratios 30-40x better than naïve approaches, enabling precise tracking of complexity dynamics. Finally, we introduce a regularization method based on spectral entropy that encourages networks toward low-complexity representations by penalizing their intrinsic dimension.

复杂度相变泛化机制正则化深度学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。