arXiv:2506.01339cs.LG2025-06ICML被引 22

让大模型遗忘更牢靠,抗住后续微调干扰

Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning

  • 引入不变性机制,通过正则化增强遗忘鲁棒性
  • 在多个任务上遗忘效果优于现有方法,且保持微调性能
  • 适合需长期保护隐私的AI系统开发者使用

机器遗忘为大语言模型的隐私与安全问题提供了潜在解决方案,可在保留模型效用的同时选择性删除特定知识。然而,现有方法对下游微调高度敏感,即使来自无关任务的微调也可能快速恢复被遗忘信息。为此,我们首次将不变性引入遗忘过程,受不变风险最小化(IRM)启发,提出基于正则化的不变式大模型遗忘框架(ILU)。ILU仅需单一数据集训练即可在多种微调任务中表现良好。通过任务向量分析进一步揭示其有效性原理。在WMDP和MUSE基准上的大量实验表明,ILU显著优于当前最优遗忘方法(如NPO和RMU),在数学、改写检测、情感分析等多样化下游任务中均实现更强的遗忘鲁棒性,同时维持良好的微调性能。

原文摘要 · Abstract (English)

Machine unlearning offers a promising solution to privacy and safety concerns in large language models (LLMs) by selectively removing targeted knowledge while preserving utility. However, current methods are highly sensitive to downstream fine-tuning, which can quickly recover forgotten information-even from unrelated tasks. To address this, we introduce invariance into unlearning for the first time, inspired by invariant risk minimization (IRM). Building on this principle, we propose invariant LLM unlearning (ILU), a regularization-based framework that enhances robustness. Notably, ILU generalizes well to diverse fine-tuning tasks, even when trained using a single dataset. A task vector analysis is also provided to further elucidate the rationale behind ILU's effectiveness. Extensive experiments on the WMDP and MUSE benchmark, reveal that ILU significantly outperforms state-of-the-art unlearning methods, including negative preference optimization (NPO) and representation misdirection for unlearning (RMU). Notably, ILU achieves superior unlearning robustness across diverse downstream fine-tuning scenarios (e.g., math, paraphrase detection, and sentiment analysis) while preserving the fine-tuning performance.

大模型遗忘隐私保护不变性微调鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。