arXiv:2607.29503cs.LG2026-07

高熵模型更抗遗忘,突破看似完美的虚假表现

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting

论文配图:The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting
图 1 · 摘自论文原文
  • 用王-兰道分子动力学采样高熵模型,对比优化性能
  • 高熵模型遗忘后仍保持95%准确率,而普通模型跌至75%以下
  • 适合关注模型鲁棒性与长期学习的研究者

尽管神经网络通常以训练和测试性能评估,但这些指标无法揭示学习表征的鲁棒性。研究表明,参数空间中占据更大体积(由玻尔兹曼熵衡量)的解往往具有更优泛化能力,即高熵优势。本文探究该优势是否超越泛化能力——即模型在后续学习新知识时能否保留旧知识。以模运算中的grokking为受控实验场景,通过噪声注入实验对比AdamW训练的Transformer与相同饱和性能下从王-兰道分子动力学采样获得的高熵模型。强制两者完全记忆带随机标签的新数据后,发现AdamW模型出现灾难性遗忘,原任务准确率从100%降至75%以下,而高熵模型保持约95%准确率。我们称此现象为‘grokked illusion’。通过奇异值分解发现,高熵模型在注意力和MLP层中均具有显著更高的有效秩,表明更丰富的特征表示可作为抗遗忘缓冲。结果表明,完美泛化不等于同等鲁棒性,为理解模型抗干扰能力提供了新视角。

原文摘要 · Abstract (English)

While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown that solutions occupying larger volumes in parameter space, as quantified by Boltzmann entropy, often exhibit superior generalizability compared to those reached by conventional optimization, a phenomenon known as the high entropy advantage. Here we ask whether this advantage persists beyond generalization. Specifically, we investigate models' robustness, the ability to retain the learned knowledge when the model is subsequently trained to acquire new information. Using grokking in modular arithmetic as a controlled setting, we design a noise injection experiment to evaluate the robustness difference between AdamW-trained transformers and high-entropy model sampled from Wang-Landau Molecular Dynamics with identical saturated performance. By forcing both models to fully remember new data with random labels, we find that AdamW-trained models suffer from catastrophic forgetting, with original task test accuracy dropping from 100% to below 75%, whereas the high-entropy models maintain approximately 95% test accuracy. We term this hidden fragility behind apparent generalization the "grokked illusion." Through singular value decomposition of the neural network weights, we discover that high-entropy neural networks possess significantly higher effective rank in attention and MLP layers both before and after noise injection, indicating richer feature representations can serve as a buffer against catastrophic forgetting. Our findings demonstrate that perfect generalization does not imply equal robustness, offering a new perspective on what makes a trained model robust to interference.

模型鲁棒性灾难性遗忘高熵优势神经网络机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。