arXiv:2502.01739cs.LGcond-mat.dis-nn2025-02被引 2

对比慢速与快速学习,发现模型特征相同但编码效率不同。

Grokking vs. Learning: Same Features, Different Encodings

  • 比较了慢速训练与快速学习的模型特征与编码效率
  • 稳态训练中出现25倍压缩率新阶段,远超慢速学习
  • 首次揭示慢速学习路径在信息空间中走直线路径

Grokking 通常能达到与普通稳定训练相当的损失水平。我们探讨了这两种不同学习路径——Grokking 与常规训练——是否导致模型本质差异。通过两个任务对比分析模型所学特征、可压缩性及学习动态,发现两者学习到的特征一致,但在特征编码效率上存在显著差异。尤其在稳态训练中,出现一种新型“压缩态”,表现为损失与可压缩性之间存在线性权衡,而Grokking中不存在此现象。在此状态下,压缩倍数可达基础模型的25倍,是Grokking压缩效果的5倍。进一步追踪训练过程,发现Grokking的模型发展具有任务依赖性,且可压缩性峰值出现在突变平台期之后。此外,引入新的信息几何度量,表明经历Grokking的模型在信息空间中沿直线路径演化。

原文摘要 · Abstract (English)

Grokking typically achieves similar loss to ordinary, "steady", learning. We ask whether these different learning paths - grokking versus ordinary training - lead to fundamental differences in the learned models. To do so we compare the features, compressibility, and learning dynamics of models trained via each path in two tasks. We find that grokked and steadily trained models learn the same features, but there can be large differences in the efficiency with which these features are encoded. In particular, we find a novel "compressive regime" of steady training in which there emerges a linear trade-off between model loss and compressibility, and which is absent in grokking. In this regime, we can achieve compression factors 25x times the base model, and 5x times the compression achieved in grokking. We then track how model features and compressibility develop through training. We show that model development in grokking is task-dependent, and that peak compressibility is achieved immediately after the grokking plateau. Finally, novel information-geometric measures are introduced which demonstrate that models undergoing grokking follow a straight path in information space.

模型压缩学习机制信息几何深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。