揭示权重衰减下模型延迟泛化的数学机制,解释为何训练后期才突然变好。
Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

- 基于线性模型与重力优化,发现延迟泛化由特定子空间主导的缓慢衰减过程。
- 理论预测的泛化延迟时间与实验观测严格吻合,且无需调参。
- 适用于理解深度学习中长期训练现象,适合研究优化机制的学者。
尽管已有大量实证研究,延迟泛化(grokking)的机理仍不清晰。本文在全批量重力优化与权重衰减的线性模型中,揭示了一种可精确求解的晚期松弛机制,并将其扩展至非线性神经网络。分析发现,经验零空间中存在一个关键的“grokking子空间”,在此子空间上训练预测保持不变,仅由权重衰减驱动缓慢耗散弛豫,遵循离散与连续时间下的精确规律。我们证明只有该子空间贡献于群体风险的渐近缓慢下降,并推导出泛化延迟时间的显式迭代尺度表达式,在弱正则化区恢复经典 $(1-β)/(ηλ)$ 标度。理论还区分了耦合 $L_2$ 正则与解耦权重衰减的影响,给出可验证的因果干预预测。我们在一个合成模型中无参数验证所有理论恒等式;并在模加法任务中观察到真实的延迟泛化,测量延迟符合预测标度,晚期弛豫与理论时钟高度一致。
原文摘要 · Abstract (English)
Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the grokking subspace. Along this subspace, the training predictions remain unchanged, leaving weight decay as the sole restoring force and giving rise to a slow dissipative relaxation governed by an exact discrete-time and continuous-time law. We show that only this subspace contributes to the slow asymptotic decay of the population risk and derive explicit iteration-scale predictions for the grokking time, recovering the familiar $(1-β)/(ηλ)$ scaling in the weak-regularization regime. The theory further predicts distinct effects of optimizer choice, distinguishing coupled $L_2$ regularization from decoupled weight decay, and yields causal predictions for interventions that modify the grokking component. We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. We further observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling and the late-time relaxation agrees closely with the theoretical clock.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。