arXiv:2411.04388cs.LG2024-11中稿 · NeurIPS被引 5

研究大模型如何删去特定数据,发现离群数据更易删除且效果更好。

Unlearning in- vs. out-of-distribution data in LLMs under gradient-based method

  • 用新指标评估大模型删忆效果,对比分布内外数据的删忆差异。
  • 删去分布外数据需更多步骤,但整体删忆效率更高。
  • 分布内数据删忆会快速降低模型性能,适合关注数据安全的研究者。

机器删忆旨在消除模型中特定训练样本的影响。尽管该问题日益受到关注,但如何评估大语言模型中的删忆效果,以及哪些数据特性会影响删忆的质量与效率,仍是未解难题。本文提出一种评估生成模型删忆质量的量化指标,并据此分析删忆质量与模型性能间的权衡。结果表明,删去分布外样本需要更多删忆步骤,但整体表现更优;而删去分布内样本时,模型性能随删忆过程迅速下降。此外,我们进一步评估了样本记忆强度与难度对经典梯度上升方法下删忆效果的影响。

原文摘要 · Abstract (English)

Machine unlearning aims to solve the problem of removing the influence of selected training examples from a learned model. Despite the increasing attention to this problem, it remains an open research question how to evaluate unlearning in large language models (LLMs), and what are the critical properties of the data to be unlearned that affect the quality and efficiency of unlearning. This work formalizes a metric to evaluate unlearning quality in generative models, and uses it to assess the trade-offs between unlearning quality and performance. We demonstrate that unlearning out-of-distribution examples requires more unlearning steps but overall presents a better trade-off overall. For in-distribution examples, however, we observe a rapid decay in performance as unlearning progresses. We further evaluate how example's memorization and difficulty affect unlearning under a classical gradient ascent-based approach.

大模型删忆性能权衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。