通过机制定位实现更鲁棒的模型知识删除与编辑
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
- 基于可解释性定位事实回忆的查找表机制进行编辑
- 在多数据集上实现更强的抗重学能力与更低副作用
- 适合需要安全可控知识管理的研究者使用
大型语言模型的知识编辑与删除方法旨在不损害通用语言建模性能的前提下,修改或移除不良知识或能力。本文研究机制可解释性(识别构成模型能力的具体机制组件)如何提升编辑与删除的精度与效果。发现不同定位方法在删除与编辑鲁棒性上存在显著差异。关键区别在于:以保持输出为主要目标的方法,与能发现具有可预测中间状态的高层机制的方法之间存在本质不同。特别是,将编辑/删除定位到事实回忆的查找表机制时,1)在多种输入输出格式下均表现出更强的鲁棒性;2)能有效抵抗重新学习被删除信息的尝试,且相比基线方法减少意外副作用,该结果在体育事实数据集和CounterFact数据集上,对多个模型均成立。此外,此类局部化编辑会比其他基线更深刻地破坏模型中的潜在知识,从而增强删除对抗各类攻击的鲁棒性。
原文摘要 · Abstract (English)
Methods for knowledge editing and unlearning in large language models seek to edit or remove undesirable knowledge or capabilities without compromising general language modeling performance. This work investigates how mechanistic interpretability -- which, in part, aims to identify model components (circuits) associated to specific interpretable mechanisms that make up a model capability -- can improve the precision and effectiveness of editing and unlearning. We find a stark difference in unlearning and edit robustness when training components localized by different methods. We highlight an important distinction between methods that localize components based primarily on preserving outputs, and those finding high level mechanisms with predictable intermediate states. In particular, localizing edits/unlearning to components associated with the lookup-table mechanism for factual recall 1) leads to more robust edits/unlearning across different input/output formats, and 2) resists attempts to relearn the unwanted information, while also reducing unintended side effects compared to baselines, on both a sports facts dataset and the CounterFact dataset across multiple models. We also find that certain localized edits disrupt the latent knowledge in the model more than any other baselines, making unlearning more robust to various attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。