arXiv:2504.12700hep-thcond-mat.dis-nn2025-04被引 4

深度学习分两个阶段:快速拟合后缓慢压缩,后者对泛化至关重要。

A Two-Phase Perspective on Deep Learning Dynamics

  • 提出深度学习双阶段模型:快速拟合与慢速压缩。
  • 三类现象均显示泛化延迟出现,且时间尺度一致。
  • 建议改进训练算法以加速压缩阶段,提升泛化能力。

我们提出,深度神经网络的学习过程分为两个阶段:快速曲线拟合阶段和缓慢的压缩或粗粒化阶段。这一观点得到三种现象的共同时间结构支持:grokking、双下降以及信息瓶颈,它们均表现出在训练误差归零后较长时间才出现泛化。我们在两种不同设置下实证验证了相关时间尺度的一致性。隐藏层与输入数据之间的互信息成为自然的进展度量,补充了基于电路的指标如局部复杂度和线性映射数。我们认为,第二阶段并非标准训练算法主动优化的部分,可能被不必要地延长。借鉴重整化群类比,我们指出该压缩阶段体现了一种有原则的遗忘形式,对泛化至关重要。

原文摘要 · Abstract (English)

We propose that learning in deep neural networks proceeds in two phases: a rapid curve fitting phase followed by a slower compression or coarse graining phase. This view is supported by the shared temporal structure of three phenomena: grokking, double descent and the information bottleneck, all of which exhibit a delayed onset of generalization well after training error reaches zero. We empirically show that the associated timescales align in two rather different settings. Mutual information between hidden layers and input data emerges as a natural progress measure, complementing circuit-based metrics such as local complexity and the linear mapping number. We argue that the second phase is not actively optimized by standard training algorithms and may be unnecessarily prolonged. Drawing on an analogy with the renormalization group, we suggest that this compression phase reflects a principled form of forgetting, critical for generalization.

深度学习泛化信息瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。