用奇异学习理论解释模型为何突然从记忆转为泛化
A Basin-Selection Perspective on Grokking via Singular Learning Theory
- 基于奇异学习理论,用局部学习系数衡量解空间盆地的统计偏好
- 低局部学习系数的盆地更易被选中,对应更低泛化误差
- 适合研究深度学习泛化机制与优化路径的科研人员
Grokking——即长时间训练后模型从记忆转向泛化的突变现象——暗示存在具有不同统计特性的竞争解盆地。本文通过奇异学习理论(SLT)框架研究该现象,该理论刻画损失曲面的几何结构。核心指标为局部学习系数(LLC),用于量化损失表面的局部退化程度。SLT表明,低LLC的盆地具有更高的后验质量集中度和更低的期望泛化误差。基于此,我们提出一种盆地选择视角:在二次网络中,LLC按统计偏好对近零损失盆地排序,而训练过程中的转换由优化动态决定。Grokking本质上是模型从高LLC(记忆型)盆地迁移到低LLC(泛化型)主导后验的盆地。我们推导了浅层二次网络在懒惰学习与特征学习两种情形下的LLC解析公式。实证上,我们展示从训练数据估计的LLC轨迹能准确追踪泛化出现的时刻,并作为优化路径的有用探针。
原文摘要 · Abstract (English)
Grokking, the abrupt transition from memorization to generalisation after extended training, suggests the presence of competing solution basins with distinct statistical properties. We study this phenomenon through the lens of Singular Learning Theory (SLT), a Bayesian framework that characterizes the geometry of the loss landscape. The key measure is the local learning coefficient (LLC) which quantifies the local degeneracy of the loss surface. SLT links lower-LLC basins to higher posterior mass concentration and lower expected generalisation error. Leveraging SLT, we develop a basin-selection perspective on grokking in quadratic networks: LLC ranks competing near-zero-loss basins by statistical preference, while the training-time transition between them is governed by optimisation dynamics. In this view, grokking corresponds to a transition from a higher-LLC (memorising) basin to a lower-LLC (generalising) basin that dominates the posterior. To support this, we derive analytic formulas for the LLC in shallow quadratic networks under both lazy and feature learning regimes. Empirically, we demonstrate that LLC trajectories estimated from training data track the onset of generalisation and provide an informative probe of the optimisation path.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。