揭示大模型知识编辑的表面性陷阱,发现旧知识仍藏在注意力机制中。
Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing
- 通过机制分析发现,编辑后旧知识仍残留在早期残差流和后期注意力模块中。
- 特定注意力头及其输出矩阵的左奇异向量与虚假编辑存在因果关系。
- 方法可推广至知识遗忘任务,对模型可解释性研究有重要参考价值。
知识编辑旨在更新语言模型中的知识,但存在欺骗性问题:尽管现有编辑算法在传统指标上表现接近完美,模型仍会生成原始知识。本文提出“表面编辑”概念描述此现象。全面评估表明该问题对现有算法构成重大挑战。系统性研究识别并验证两个关键因素:(1) 较早层中最后主题位置的残差流;(2) 较晚层中的特定注意力模块。值得注意的是,后期某些注意力头及其输出矩阵的特定左奇异向量封装了原始知识,并与表面编辑存在因果关系。此外,我们将分析扩展至表面去学习任务,观察到相同注意力头及对应左奇异向量的行为模式,证明了方法与结论的稳健性和广泛适用性。代码已公开。
原文摘要 · Abstract (English)
Knowledge editing, which aims to update the knowledge encoded in language models, can be deceptive. Despite the fact that many existing knowledge editing algorithms achieve near-perfect performance on conventional metrics, the models edited by them are still prone to generating original knowledge. This paper introduces the concept of "superficial editing" to describe this phenomenon. Our comprehensive evaluation reveals that this issue presents a significant challenge to existing algorithms. Through systematic investigation, we identify and validate two key factors contributing to this issue: (1) the residual stream at the last subject position in earlier layers and (2) specific attention modules in later layers. Notably, certain attention heads in later layers, along with specific left singular vectors in their output matrices, encapsulate the original knowledge and exhibit a causal relationship with superficial editing. Furthermore, we extend our analysis to the task of superficial unlearning, where we observe consistent patterns in the behavior of specific attention heads and their corresponding left singular vectors, thereby demonstrating the robustness and broader applicability of our methodology and conclusions. Our code is available here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。