提出新方法让神经网络跳过死记硬背,更快学会通用规律。
Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping

- 用参数移动性、副本相关性等指标量化训练过程中的玻璃态特征。
- 发现标准优化会困在记忆陷阱中,参数移动性崩溃且依赖历史。
- 设计蒙特卡洛参数交换机制,模拟物理扩散加速泛化,效果优于常规方法。
Grokking 是神经网络训练中一种显著现象:模型先经历长期纯记忆阶段,随后突然实现泛化。尽管已有研究将其类比为统计物理中的玻璃弛豫,将初期记忆视为快速冷却形成的玻璃态,后期泛化视为缓慢弛豫,但这一理论仍停留在宏观层面,缺乏对训练动态的直接实证。本文提出三组件框架,通过参数移动性(PM)、副本相关性(RC)和分形维度(FD)量化训练过程中的玻璃动力学特征。结果表明,标准优化呈现出明显的玻璃态行为,使模型陷入动能冻结的记忆状态,表现为参数移动性塌陷、强历史依赖性和通道式运动。基于此,我们提出状态感知的蒙特卡洛参数交换(SAM-Swap),受玻璃动力学中交换蒙特卡洛算法启发。对比 SAM-Swap、权重衰减和高斯梯度噪声,发现加速泛化的共同特征是参数空间中的随机探索,类似于物理中的扩散过程。
原文摘要 · Abstract (English)
Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works have attempted to interpret it through classical machine learning mechanisms like weight norm, recent research draws an analogy from statistical physics, framing grokking as a form of computational glass relaxation. This theory defines the initial memorization as a result of `fast cooling' where the training loss is reduced so quickly that a glass state is formed, followed by a `slow relaxation' towards final generalization. Although providing a unifying framework for representative grokking theories, this perspective has remained largely at the theoretical on macroscopic level without direct empirical validation on training dynamics. Here we introduce a three-component framework to directly characterize the training dynamics via parameter mobility (PM), and two representative measurements from glassy dynamics: replica correlation (RC) and fractal dimension (FD). We demonstrate that standard optimization presents clear signatures of glass dynamics and inherently traps the grokking network in a kinetic arrested memorization state with a collapsed mobility, strong history dependence, and channel-like motions. This quantitative agreement motivates us to introduce State-Aware Monte Carlo Parameter Swapping (SAM-Swap), an optimization plug-in that can accelerate generalization, inspired by swap Monte Carlo algorithm widely used in glass dynamics. Comparing SAM-Swap, weight decay, and Gaussian gradient noise, we find that accelerated generalization is consistently associated with random exploration in the parameter space, similar to diffusion in physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。