arXiv:2505.19855cs.LG2025-05AAAI被引 14

知识编辑方法可作大模型删忆强基线,尤其适合预训练知识删除。

Editing as Unlearning: Are Knowledge Editing Methods Strong Baselines for Large Language Model Unlearning?

  • 将删忆视为特殊编辑,把信息改为空拒绝响应
  • WISE和AlphaEdit在预训练知识删忆上效果出色
  • 提出自提升与查询合并策略,提升编辑适配性

大语言模型(LLM)删忆对负责任部署至关重要。不同于知识编辑仅修改知识,删忆需选择性移除信息。本文发现两者紧密关联:删忆可视为一种特殊编辑,即将信息修改为拒绝或空集∅响应。我们评估了主流编辑方法(如ROME、MEMIT、GRACE、WISE、AlphaEdit)在预训练与微调知识上的删忆表现。结果表明,特别是WISE和AlphaEdit,在预训练知识删忆中表现优异,且生成更符合人类对齐的拒绝回答。为更好适配编辑方法用于删忆,我们提出实用方案:自提升利用模型自身上下文学习能力构建更对齐的删忆目标;查询合并使ROME和MEMIT能有效处理长序列样本。建议删忆领域采用先进编辑方法作为基线,从编辑视角探索更全面的模型记忆控制。

原文摘要 · Abstract (English)

Large language Model (LLM) unlearning, i.e., selectively removing information from LLMs, is vital for responsible model deployment. Differently, LLM knowledge editing aims to modify LLM knowledge instead of removing it. Though editing and unlearning seem to be two distinct tasks, we find there is a tight connection between them. In this paper, we conceptualize unlearning as a special case of editing where information is modified to a refusal or "empty set" $\emptyset$ response, signifying its removal. This paper thus investigates if knowledge editing techniques are strong baselines for LLM unlearning. We evaluate state-of-the-art (SOTA) editing methods (e.g., ROME, MEMIT, GRACE, WISE, and AlphaEdit) against existing unlearning approaches on pretrained and finetuned knowledge. Results show certain editing methods, notably WISE and AlphaEdit, are effective unlearning baselines, especially for pretrained knowledge, and excel in generating human-aligned refusal answers. To better adapt editing methods for unlearning applications, we propose practical recipes including self-improvement and query merging. The former leverages the LLM's own in-context learning ability to craft a more human-aligned unlearning target, and the latter enables ROME and MEMIT to perform well in unlearning longer sample sequences. We advocate for the unlearning community to adopt SOTA editing methods as baselines and explore unlearning from an editing perspective for more holistic LLM memory control.

删忆知识编辑大模型基线

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。