解释了模型延迟泛化现象的几何机制,揭示了学习率等参数如何影响泛化出现时间。
A Stochastic--Geometric Theory of Scaling Laws in Grokking
- 基于Adam优化器与权重衰减,发现参数空间中存在壳层-核心拓扑结构。
- 推导出学习率、批量大小和正则化系数的泛化延迟缩放规律。
- 适合研究深度学习优化机制或泛化现象的学者参考。
延迟泛化(即“领悟”)指神经网络在训练初期快速拟合数据,但泛化能力在长时间后才突然出现的现象。尽管已有大量实证研究,其内在机制仍不清晰。本文首次从理论上刻画了由Adam优化动态与权重衰减正则化所诱导的可解空间的壳-核拓扑结构,并通过实证支持。在参数空间中,随机初始化的解集中在薄外球壳,包围着记忆解的球壳,其内部为泛化解的核心区域。利用停止时间理论,分析该拓扑结构下的解路径转移时间,即优化轨迹何时脱离记忆流形并首次抵达泛化解流形边界。理论推导出学习率、批量大小和ℓ₂正则化系数的领悟缩放规律,实验验证并与先前文献结果一致。
原文摘要 · Abstract (English)
Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition. Despite extensive empirical study, its underlying mechanism remains poorly understood. In this work, we first theoretically characterize a shell--core topological configuration of the reachable solution space induced by Adam's optimization dynamics with weight-shrinkage regularization, supported by empirical evidence. This optimization-induced topological configuration gives rise to grokking. In model's parameter space, random initialization solutions concentrate on a thin outer spherical shell, enclosing another spherical shell of memorization solutions, which in turn contains a core corresponding to the generalization solutions. Leveraging stopping-time theory, we then analyze the geometry of this topological configuration and the solution transition time at which optimization trajectories escape the memorization manifold and first reach the boundary of the generalization manifold. Our theoretical analysis derives grokking scaling laws for the learning rate, batch size, and $\ell_2$ regularization coefficient, which are further validated through experiments and shown to recover results from prior literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。