发现深度模型训练中突然从记不住到完美泛化的现象及其机制
Exploring Grokking: Experimental and Mechanistic Investigations
- 通过大量实验研究模型、数据量与优化对突变泛化的影响
- 训练初期完全记忆数据,长期训练后出现泛化能力的急剧跃迁
- 适合关注神经网络内在机制与训练动力学的研究者
过参数化神经网络中的grokking现象引起广泛关注。该现象表现为:模型在训练初期能将训练集完全记忆(训练误差为零),但测试误差接近随机水平;经过长时间持续训练后,泛化能力出现剧烈跃迁,实现完美泛化。本研究包含大规模实验及对grokking机制的深入探讨。我们通过实验揭示了其在不同训练数据比例、模型结构和优化过程下的行为特征,并梳理了学术界关于该现象的多种解释视角。
原文摘要 · Abstract (English)
The phenomenon of grokking in over-parameterized neural networks has garnered significant interest. It involves the neural network initially memorizing the training set with zero training error and near-random test error. Subsequent prolonged training leads to a sharp transition from no generalization to perfect generalization. Our study comprises extensive experiments and an exploration of the research behind the mechanism of grokking. Through experiments, we gained insights into its behavior concerning the training data fraction, the model, and the optimization. The mechanism of grokking has been a subject of various viewpoints proposed by researchers, and we introduce some of these perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。