通过曲率分解揭示模型记忆机制,有效清除无关记忆而不损核心能力
From Memorization to Reasoning in the Spectrum of Loss Curvature
- 基于损失曲率高低排序权重,分离出记忆相关成分
- 删除低曲率权重后,非目标记忆减少超平衡子网法,困惑度更低
- 数学与事实检索性能下降,证明任务依赖特定权重方向
我们刻画了变压器模型中记忆的表征方式,发现可通过损失曲率谱对语言模型(LMs)和视觉变压器(ViTs)的权重进行解耦。该方法基于理论与实证工作:记忆数据对应的曲率远高于非记忆数据,因此按曲率从高到低排序权重可无标签区分记忆成分。这启发了一种权重编辑方法,相比近期的去记忆方法(BalancedSubnet),能更有效地抑制非目标记忆,同时保持更低的困惑度。由于曲率基具有权重共享结构的自然解释,我们在多个下游任务上系统分析了该编辑的影响,发现事实检索与算术任务性能显著下降,而开卷事实检索与一般逻辑推理能力得以保留。我们认为这些任务高度依赖权重空间中的特殊方向,而非通用机制,无论具体数据是否被记忆。我们进一步验证:任务数据激活强度与被删除的低曲率组件呈正相关,且任务性能下降与此对应。本研究深化了对神经网络记忆的理解,并为实际去除记忆提供了新路径,同时揭示了数学与事实检索等任务所依赖的独特、窄范围结构。
原文摘要 · Abstract (English)
We characterize how memorization is represented in transformer models and show that it can be disentangled in the weights of both language models (LMs) and vision transformers (ViTs) using a decomposition based on the loss landscape curvature. This insight is based on prior theoretical and empirical work showing that the curvature for memorized training points is much sharper than non memorized, meaning ordering weight components from high to low curvature can reveal a distinction without explicit labels. This motivates a weight editing procedure that suppresses far more recitation of untargeted memorized data more effectively than a recent unlearning method (BalancedSubnet), while maintaining lower perplexity. Since the basis of curvature has a natural interpretation for shared structure in model weights, we analyze the editing procedure extensively on its effect on downstream tasks in LMs, and find that fact retrieval and arithmetic are specifically and consistently negatively affected, even though open book fact retrieval and general logical reasoning is conserved. We posit these tasks rely heavily on specialized directions in weight space rather than general purpose mechanisms, regardless of whether those individual datapoints are memorized. We support this by showing a correspondence between task data's activation strength with low curvature components that we edit out, and the drop in task performance after the edit. Our work enhances the understanding of memorization in neural networks with practical applications towards removing it, and provides evidence for idiosyncratic, narrowly-used structures involved in solving tasks like math and fact retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。