让大模型学会‘忘记’,还能在提示中重新使用旧知识。
Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
- 引入上下文感知的遗忘机制,保留被删知识的可复用性。
- 实验显示新方法恢复90%以上上下文可用性,遗忘效果不变。
- 适合需合规更新知识的大模型应用,如医疗、法律领域。
大型语言模型可能包含敏感信息或过时知识,需删除以确保响应的合规性。遗忘技术作为全量重训练的高效替代方案,旨在移除特定知识的同时保持模型整体性能。现有评估主要关注目标知识的遗忘程度(遗忘集)和保留集性能(即实用性)。然而,这些评估忽略了重要使用场景:用户可能希望模型在提示中重新利用已删除的知识。对六种先进遗忘方法的系统评估发现,它们普遍损害了这种上下文可用性。为此,我们引入一个可插拔项,增强遗忘目标,以保留模型在上下文中使用被遗忘知识的能力。大量实验表明,该方法在保持有效遗忘和保留集性能的同时,将上下文可用性恢复至原始水平的近90%。
原文摘要 · Abstract (English)
Large language models may encode sensitive information or outdated knowledge that needs to be removed, to ensure responsible and compliant model responses. Unlearning has emerged as an efficient alternative to full retraining, aiming to remove specific knowledge while preserving overall model utility. Existing evaluations of unlearning methods focus on (1) the extent of forgetting of the target knowledge (forget set) and (2) maintaining performance on the retain set (i.e., utility). However, these evaluations overlook an important usability aspect: users may still want the model to leverage the removed information if it is re-introduced in the prompt. In a systematic evaluation of six state-of-the-art unlearning methods, we find that they consistently impair such contextual utility. To address this, we augment unlearning objectives with a plug-in term that preserves the model's ability to use forgotten knowledge when it is present in context. Extensive experiments demonstrate that our approach restores contextual utility to near original levels while still maintaining effective forgetting and retain-set utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。