通过梯度比定位关键参数,精准清除大模型记忆而不伤及其他知识。
LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
- 用梯度比衡量参数对遗忘数据的贡献,精确定位需处理的权重。
- 对关键参数注入噪声后微调,实现90%以上记忆清除且性能下降小于5%。
- 专为大模型设计评估指标,适用于隐私保护与合规场景。
随着大型语言模型(LLM)面临日益增长的法律与伦理审查,有效的机器遗忘机制变得至关重要,尤其是针对敏感或未经授权的数据。现有经验方法常导致遗忘不彻底或无关知识意外退化,根源在于定位不准确。本文提出GRIN:一种模块化、目标明确的LLM遗忘框架。GRIN引入基于梯度比的新度量方法,识别对记忆遗忘数据贡献最大的参数。随后在微调前对这些参数进行选择性噪声注入,提升遗忘效果同时保持模型实用性。最后,我们提出了针对LLM场景的新评估指标,并在TOFU、WMDP和SafePKU等标准基准上验证了该方法的有效性。
原文摘要 · Abstract (English)
The growing legal and ethical scrutiny of large language models (LLMs) necessitates effective machine unlearning, particularly for sensitive or unauthorized data. Existing empirical methods often yield incomplete forgetting or unintended degradation of unrelated knowledge due to poor localization. In this work, we propose GRIN: a modular and targeted framework for LLM unlearning. GRIN introduces a novel gradient-ratio-based metric to identify parameters most responsible for memorizing forget data. We then perform selective noise injection into these parameters prior to fine-tuning, which improves unlearning performance while maintaining model utility. Finally, we propose new evaluation metrics tailored to the LLM setting and validate our approach on standard benchmarks such as TOFU, WMDP, and SafePKU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。