arXiv:2601.19791cs.LGstat.ML2026-01中稿 · ICML被引 3

揭示线性回归中泛化延迟的数学机制,证明可通过调参控制过拟合后的突然泛化。

To Grok Grokking: Provable Grokking in Ridge Regression

  • 在带权重衰减的梯度下降中,证明了过拟合后长期不泛化、最终突然泛化的完整过程。
  • 首次给出泛化延迟时间(grokking time)与训练超参数的严格定量关系。
  • 结果表明过拟合后的突然泛化是可调控的训练现象,非模型本质缺陷。

我们研究经典岭回归设置下的grokking现象——即在过拟合之后很久才出现泛化能力。通过梯度下降配合权重衰减,我们证明了过参数化线性回归模型的端到端grokking过程:(i) 模型在训练初期即过拟合训练数据;(ii) 过拟合后泛化性能持续较差;(iii) 最终泛化误差趋于任意小。此外,理论与实证均表明,通过合理调整超参数可放大或消除grokking。据我们所知,这是首个关于泛化延迟时间(称作“grokking time”)与训练超参数之间严格定量边界的成果。最后,我们还通过实验发现,这些定量边界在非线性神经网络中同样适用。结果表明,grokking并非深度学习的固有失败模式,而是特定训练条件的结果,无需修改模型架构或学习算法即可避免。

原文摘要 · Abstract (English)

We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting. We prove end-to-end grokking results for learning over-parameterized linear regression models using gradient descent with weight decay. Specifically, we prove that the following stages occur: (i) the model overfits the training data early during training; (ii) poor generalization persists long after overfitting has manifested; and (iii) the generalization error eventually becomes arbitrarily small. Moreover, we show, both theoretically and empirically, that grokking can be amplified or eliminated in a principled manner through proper hyperparameter tuning. To the best of our knowledge, these are the first rigorous quantitative bounds on the generalization delay (which we refer to as the "grokking time") in terms of training hyperparameters. Lastly, going beyond the linear setting, we empirically demonstrate that our quantitative bounds also capture the behavior of grokking on non-linear neural networks. Our results suggest that grokking is not an inherent failure mode of deep learning, but rather a consequence of specific training conditions, and thus does not require fundamental changes to the model architecture or learning algorithm to avoid.

泛化延迟岭回归梯度下降超参数调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。