模型容量决定突现学习,靠记忆与泛化速度的较量。
Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds

- 用记忆速度和泛化速度的竞赛解释突现学习现象。
- 当参数量使两速度相交时,突现学习开始发生。
- 适合研究模型容量与学习行为关系的学者参考。
现有对突现学习的解释多基于机制框架,如电路效率或懒惰到丰富过渡。尽管已知突现学习与模型大小相关,但模型容量如何影响这一现象仍不明确。本文针对模运算任务,提出信息论视角:突现学习并非在模型大到能记忆训练集时立即出现,而是记忆速度 $T_{\text{mem}}(P)$ 与泛化速度 $T_{\text{gen}}(P)$ 两个可测量时间尺度竞争的结果,二者均依赖于模型参数量 $P$。通过借鉴 Morris 等(2025)的信息容量框架,我们分别在等复杂度的随机标签数据上估计 $T_{\text{mem}}(P)$,在模运算任务上估计 $T_{\text{gen}}(P)$,发现突现学习出现在两时间尺度交汇处附近。该框架还提出一个经验模型,用于预测给定容量和数据复杂度下的记忆速度,复现了大模型记忆更快的已有观察。整体表明,学习时间尺度的正式化是理解模型容量如何塑造算法任务上突现学习的关键抽象。
原文摘要 · Abstract (English)
Existing accounts of grokking explain the phenomena in terms of mechanistic frameworks such as circuit efficiency or lazy-to-rich transitions. However, despite a known dependence between grokking and model size, how model capacity shapes grokking remains an open question. We give an information-theoretic account of this relationship on the task of modular arithmetic, showing that grokking does not immediately occur when a model becomes large enough to memorise the training set, but rather emerges as the outcome of a competition between two measurable timescales: a memorisation speed $T_{\text{mem}}(P)$ and a generalisation speed $T_{\text{gen}}(P)$, both of which are functions of model parameter count $P$. Adapting the information capacity framework of Morris et al. (2025), we estimate $T_{\text{mem}}(P)$ on random-label data of equivalent complexity and $T_{\text{gen}}(P)$ on the modular task itself, and show that grokking emerges close to the parameter scale where these timescales intersect. The framework also suggests an empirical model for predicting memorisation speed given model capacity and dataset complexity, recovering the previously reported empirical observation that larger models memorise faster. Overall, we motivate the formalisation of different learning timescales as important abstractions to study when explaining how model capacity shapes grokking on algorithmic tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。