揭露大模型知识编辑的虚假删除现象,揭示其本质是抑制而非消除。
Exposing the Illusion of Erasure in Knowledge Editing for LLMs

- 从对抗性触发角度分析,编辑知识实际未被彻底清除。
- 低秩更新仅重分布知识,不覆盖原有信息,原知识仍可被激发。
- 编辑内容位于敏感区域,易受提示扰动和攻击,适合安全研究者关注。
知识编辑(KE)作为无需昂贵重训练即可更新大语言模型特定事实的前沿技术,其可靠性和内在机制仍不明确。本文从对抗性触发视角审视KE,发现编辑后的知识常未被完全擦除,且在多种模型架构中均存在一致失败。通过机制分析,我们发现低秩更新并非覆盖原有知识,而是将其重新分布于模型表示空间中。此外,这些方法实为定向抑制机制,降低原始事实的表达概率,而非将其移除。损失景观分析表明,编辑知识位于狭窄、各向异性的区域,对扰动高度敏感,极易受间接提示和对抗攻击影响。本研究揭示了架构层面的根本漏洞,证明现有KE算法本质上可被绕过,呼吁对后训练更新在多个应用场景中的部署进行根本性反思。
原文摘要 · Abstract (English)
Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an adversarial elicitation perspective, revealing that edited knowledge is often not fully erased and continues to surface, with consistent failures observed across diverse model architectures. To explain this behavior, we conduct a mechanistic analysis of popular KE methods. We show that low-rank updates do not overwrite existing knowledge but instead redistribute it within the model's representation space. Furthermore, we find that these methods act as targeted suppression mechanisms that reduce the likelihood of expressing original facts, rather than removing them from the model. Analysis of the loss landscape reveals that edited knowledge lies in narrow, anisotropic regions that are highly sensitive to perturbations, making them highly vulnerable to indirect prompting and adversarial attacks. By exposing these profound architectural vulnerabilities, our work proves that KE algorithms are inherently bypassable and motivates a fundamental reevaluation of how we deploy post-hoc updates in several LLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。