arXiv:2606.05863cs.LGcs.AI2026-06被引 1

揭示模型训练中拟合与简化表示的双重时钟现象。

Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction

论文配图:Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction
图 1 · 摘自论文原文
  • 用线性网络理论分离损失下降与表征简化的时间尺度
  • 线性模型中损失在对数时间收敛,非线性时呈现多项式收敛
  • 适用于理解深度神经网络的两阶段学习机制

Grokking 指出,模型对训练数据的拟合与学习简单底层规律可能发生在不同时间尺度。我们通过将分类损失的快速衰减与学习表征的缓慢简化分离,定义了两个停止时间,即‘双重训练时钟’。对于深度线性网络,后边界间隙增长或一步尾部收缩条件可在对数时间尺度上使交叉熵损失降至 ε 水平。当存在逐层权重衰减时,端到端映射的正则化可表达为施瓦茨型惩罚;在尖锐的晚期柯尔达-洛贾斯维茨尾部条件下,该结构能量在多项式时间尺度上闭合。因此,两个时钟分离开拟合与表征简化过程。我们进一步解释该机制如何在带 ReLU 的 MLP 中出现:在训练集激活模式保持不变的区域,网络退化为活跃坐标上的线性模型。在两层 ReLU 嵌入模型中,链式法则估计表明分类头可获得比嵌入块更大的有效梯度,且受控下游范数下成立。这支持两阶段机制:分类器先拟合,而表征随后持续简化。实验以模加法为主要设置。线性理论提供严格分析核心,但 ReLU 结果为条件性简化,解释经验行为而不声称对非线性训练动态的全局证明。

原文摘要 · Abstract (English)

Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales. We formalize this phenomenon by separating the fast decay of the classification loss from the slower simplification of the learned representation, and we call the resulting pair of stopping times two training clocks. For deep linear networks, we show that a post-margin gap-growth or one-step tail-contraction condition reduces the cross-entropy loss to level epsilon on a logarithmic time scale. In contrast, when layerwise weight decay is present, the induced regularization on the end-to-end map can be expressed as a Schatten-type penalty; under a sharp late-time Kurdyka-Lojasiewicz tail, this structural energy closes on a polynomial time scale. The two clocks, therefore, separate fitting from representation simplification. We then explain how the same mechanism can appear in ReLU MLPs. In regions where the activation patterns on the training set remain fixed, the network reduces to a linear model in the active coordinates. In a two-layer ReLU embedding model, chain-rule estimates further show that the classifier head can receive larger effective gradients than the embedding block under controlled downstream norms. This supports a two-stage mechanism in which the classifier fits first, while the representation continues to simplify later. We use modular addition as the main experimental setting. The deep linear theory provides the rigorous core of the analysis. But the ReLU results are formulated as conditional reductions that account for empirical behavior without claiming a global proof for nonlinear training dynamics.

深度学习训练动态表征学习线性理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。