arXiv:2502.19301cs.LG2025-02ICLR被引 66

从梯度角度重新审视大模型知识删除,揭示现有方法缺陷并提出改进方向。

Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond

  • 通过梯度效应分析不同知识删除目标的影响机制。
  • 发现现有方法在多层、多实例、多步骤中存在性能失衡问题。
  • 为大模型安全更新提供可解释的评估工具,适合模型安全研究者。

大语言模型需经过严格审计以识别潜在风险,如版权和隐私侵犯。一旦发现问题,及时更新至关重要,以移除不良响应,确保模型合法安全使用。这推动了大模型知识删除研究的发展,旨在消除特定不良知识而不损害其他非目标响应的完整性。现有研究提出了多种知识删除目标,无需完全重训练即可实现。然而,这些目标各有特性,尚无统一框架深入理解其效果。为此,我们提出梯度效应(G-effect)工具包,从梯度视角量化知识删除目标对模型性能的影响。该工具具备广泛能力,可从实例、更新步骤和模型层等多个维度详细分析删除影响。G-effect为揭示现有知识删除目标的缺陷提供了新洞见,并进一步促使我们探索一系列缓解与优化方案。最后,我们指出若干值得深入研究的未来方向,旨在推动该重要领域的发展。

原文摘要 · Abstract (English)

Large language models (LLMs) should undergo rigorous audits to identify potential risks, such as copyright and privacy infringements. Once these risks emerge, timely updates are crucial to remove undesirable responses, ensuring legal and safe model usage. It has spurred recent research into LLM unlearning, focusing on erasing targeted undesirable knowledge without compromising the integrity of other, non-targeted responses. Existing studies have introduced various unlearning objectives to pursue LLM unlearning without necessitating complete retraining. However, each of these objectives has unique properties, and no unified framework is currently available to comprehend them thoroughly. To fill the gap, we propose a toolkit of the gradient effect (G-effect), quantifying the impacts of unlearning objectives on model performance from a gradient perspective. A notable advantage is its broad ability to detail the unlearning impacts from various aspects across instances, updating steps, and LLM layers. Accordingly, the G-effect offers new insights into identifying drawbacks of existing unlearning objectives, further motivating us to explore a series of new solutions for their mitigation and improvements. Finally, we outline promising directions that merit further studies, aiming at contributing to the community to advance this important field.

大模型安全知识删除梯度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。