提出可选择性撤销模型知识修改的方法,避免误删有用信息。
Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage

- 基于奇异向量稀疏编码假设,定位编辑敏感成分。
- 在多个设置中成功逆转指定编辑,同时保留其他编辑结果。
- 适合需要精准修复知识的场景,如安全可控的模型更新。
知识编辑为高效更新大语言模型中的事实知识提供了途径。然而,恶意编辑可能引入安全风险,因此有必要逆转不良编辑效果。现有针对参数修改的逆转方法主要关注全局消除,可能一并抹除应保留的有益编辑。本文研究选择性逆转编辑知识,目标是在不破坏其他编辑的前提下,仅逆转特定事实。基于每个编辑在权重矩阵主子空间中稀疏编码的假设,我们提出一种基于谱的逆转框架,可在编辑权重的主奇异子空间中定位编辑敏感组件。在多种设置下的实验表明,该方法能有效逆转选定编辑,同时保留无关编辑内容。结果表明,不同编辑在主奇异分量中稀疏分布,且当编辑数量适中时可分离,使得选择性谱逆转成为定位编辑特异性组件、修复编辑后语言模型的有前景方向。
原文摘要 · Abstract (English)
Knowledge editing provides an efficient way to update factual knowledge in large language models. However, malicious edits may introduce safety risks, making it necessary to reverse undesirable editing effects. Existing reversal methods for parameter-modifying edits mainly focus on global removal, which may also erase beneficial edits that should be preserved. In this paper, we study selective reversal of edited knowledge, where the goal is to reverse targeted edited facts while preserving the remaining edited facts. Based on the hypothesis that each edit is sparsely encoded within the dominant subspace of the edited matrix, we propose a spectral-based reversal framework that locates edit-sensitive components within the dominant singular subspace of edited weights. Experiments across multiple settings demonstrate the effectiveness of our method in reversing selected edits while preserving unrelated edited facts. These results suggest that different edits are sparsely encoded within dominant singular components and can be separable when the number of edits is moderate, making selective spectral reversal a promising direction for locating edit-specific components and repairing edited language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。