揭示模运算中模型从记忆到泛化的结构简化机制
Grokking From Abstraction to Intelligence
- 基于简约原则,发现模型在泛化时自发简化内部结构
- 泛化转变对应冗余流形坍缩与深度信息压缩
- 适合研究模型泛化、过拟合机制的学者参考
模运算中的grokking现象已成为研究模型泛化机制的核心实验范式。现有研究多聚焦局部电路或优化调参,忽视了驱动该现象的全局结构演化。本文提出,grokking源于模型内部结构在简约原则下自发简化。通过结合因果、谱和算法复杂度度量,以及奇异学习理论,揭示从记忆到泛化的转变对应冗余流形的物理坍缩与深度信息压缩,为理解模型过拟合与泛化机制提供了新视角。
原文摘要 · Abstract (English)
Grokking in modular arithmetic has established itself as the quintessential fruit fly experiment, serving as a critical domain for investigating the mechanistic origins of model generalization. Despite its significance, existing research remains narrowly focused on specific local circuits or optimization tuning, largely overlooking the global structural evolution that fundamentally drives this phenomenon. We propose that grokking originates from a spontaneous simplification of internal model structures governed by the principle of parsimony. We integrate causal, spectral, and algorithmic complexity measures alongside Singular Learning Theory to reveal that the transition from memorization to generalization corresponds to the physical collapse of redundant manifolds and deep information compression, offering a novel perspective for understanding the mechanisms of model overfitting and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。