通过清除相关知识提升大模型删忆效果
UIPE: Enhancing LLM Unlearning by Removing Knowledge Related to Forgetting Targets
- 基于参数外推法移除与遗忘目标相关的逻辑知识
- 在TOFU基准上显著提升多种主流删忆方法性能
- 适合需要精准删除敏感信息的模型安全场景
大型语言模型在海量数据训练中不可避免地习得有害信息。删忆旨在消除这些有害信息的影响,同时保持模型整体性能。现有基于梯度上升的方法主要关注遗忘目标数据,却忽略了逻辑相关知识对删忆效果的关键影响。本文通过理论与实验分析发现,模型可通过逻辑相关知识重构目标内容,是导致删忆性能不佳的主要原因。为此,提出基于参数外推的删忆增强方法(UIPE),主动移除与遗忘目标高度相关的知识。实验表明,UIPE显著提升了多种主流大模型删忆方法在TOFU基准上的表现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) inevitably acquire harmful information during training on massive datasets. LLM unlearning aims to eliminate the influence of such harmful information while maintaining the model's overall performance. Existing unlearning methods, represented by gradient ascent-based approaches, primarily focus on forgetting target data while overlooking the crucial impact of logically related knowledge on the effectiveness of unlearning. In this paper, through both theoretical and experimental analyses, we first demonstrate that a key reason for the suboptimal unlearning performance is that models can reconstruct the target content through reasoning with logically related knowledge. To address this issue, we propose Unlearning Improvement via Parameter Extrapolation (UIPE), a method that removes knowledge highly correlated with the forgetting targets. Experimental results show that UIPE significantly enhances the performance of various mainstream LLM unlearning methods on the TOFU benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。