将神经网络的突现泛化现象类比为玻璃弛豫,揭示其非平衡态本质。
Is Grokking a Computational Glass Relaxation?
- 把神经网络参数看作物理系统自由度,训练损失为能量,模拟玻璃形成过程。
- 实验表明记忆到泛化的转变无熵屏障,挑战了相变理论的假设。
- 提出新优化器WanD,可消除突现泛化,为优化器设计提供新思路。
理解神经网络(NN)的泛化能力是深度学习的核心问题。突现泛化(grokking)这一特殊现象——即在训练误差接近完美后,模型突然实现泛化——为研究神经网络泛化机制提供了独特窗口。本文将突现泛化解释为计算玻璃弛豫:将神经网络视为一个物理系统,参数为自由度,训练损失为系统能量,发现记忆过程类似于低温下液体快速冷却至非平衡玻璃态,而后续泛化则类似缓慢弛豫至更稳定构型。该映射使我们能够采样神经网络的玻尔兹曼熵(状态密度)随训练损失和测试准确率变化的景观。在变换器模型上的算术任务实验表明,记忆到泛化的过渡过程中不存在熵屏障,挑战了此前将突现泛化视为一阶相变的理论。我们发现了突现泛化下的高熵优势,扩展了熵与泛化能力关联的研究,且效应更为显著。受突现泛化远离平衡态的启发,我们设计了一种基于王-兰道分子动力学的玩具优化器WanD,可在不加约束条件下消除突现泛化,并找到高范数的泛化解。这为反对仅由权重范数演化进入‘恰到好处’区域导致突现泛化的理论提供了严格反例,也暗示了优化器设计的新方向。
原文摘要 · Abstract (English)
Understanding neural network's (NN) generalizability remains a central question in deep learning research. The special phenomenon of grokking, where NNs abruptly generalize long after the training performance reaches a near-perfect level, offers a unique window to investigate the underlying mechanisms of NNs' generalizability. Here we propose an interpretation for grokking by framing it as a computational glass relaxation: viewing NNs as a physical system where parameters are the degrees of freedom and train loss is the system energy, we find memorization process resembles a rapid cooling of liquid into non-equilibrium glassy state at low temperature and the later generalization is like a slow relaxation towards a more stable configuration. This mapping enables us to sample NNs' Boltzmann entropy (states of density) landscape as a function of training loss and test accuracy. Our experiments in transformers on arithmetic tasks suggests that there is NO entropy barrier in the memorization-to-generalization transition of grokking, challenging previous theory that defines grokking as a first-order phase transition. We identify a high-entropy advantage under grokking, an extension of prior work linking entropy to generalizability but much more significant. Inspired by grokking's far-from-equilibrium nature, we develop a toy optimizer WanD based on Wang-landau molecular dynamics, which can eliminate grokking without any constraints and find high-norm generalizing solutions. This provides strictly-defined counterexamples to theory attributing grokking solely to weight norm evolution towards the Goldilocks zone and also suggests new potential ways for optimizer design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。