区分预训练与微调阶段知识,发现微调后更易安全删减信息
Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning
- 构建28.6万个三元组的双阶段遗忘评估基准
- 微调模型比预训练模型遗忘更平滑,保留率高10%-50%
- 揭示知识来源影响遗忘效果,适合模型安全更新研究者
机器遗忘(MU)使大语言模型能够删除不安全或过时的信息。然而,现有工作假设所有知识都同等可遗忘,且忽视了被遗忘知识源自预训练还是监督微调(SFT)。本文提出DUET(跨训练阶段双重遗忘评估),包含28.6万个源自Wikidata的三元组,通过维基链接数和LLM生成的显著性评分标注事实流行度。实验表明,预训练模型与微调模型对遗忘的响应不同:在遗忘数据上进行一次SFT可实现更平滑的遗忘、更稳定的调参,且保留率高出10%-50%;而直接对预训练模型执行遗忘仍不稳定,易出现重学或灾难性遗忘。
原文摘要 · Abstract (English)
Machine Unlearning (MU) enables Large Language Models (LLMs) to remove unsafe or outdated information. However, existing work assumes that all facts are equally forgettable and largely ignores whether the forgotten knowledge originates from pretraining or supervised fine-tuning (SFT). In this paper, we introduce DUET (Dual Unlearning Evaluation across Training Stages), a benchmark of 28.6k Wikidata-derived triplets annotated with fact popularity using Wikipedia link counts and LLM-based salience scores. Our experiments show that pretrained and SFT models respond differently to unlearning. An SFT step on the forget data yields smoother forgetting, more stable tuning, and 10-50% higher retention, while direct unlearning on pretrained models remains unstable and prone to relearning or catastrophic forgetting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。