arXiv:2410.22086cs.LGcs.CL2024-10NAACL被引 35

提出新方法让大模型高效删知识,兼顾遗忘与性能。

Unlearning as multi-task optimization: A normalized gradient difference approach with an adaptive learning rate

  • 将删知识视为多任务优化,平衡遗忘与模型能力。
  • 在TOFU和MUSE数据集上效果优于现有方法,训练稳定。
  • 自适应学习率+归一化梯度差,控制遗忘与性能平衡。

机器遗忘被用于移除大型语言模型(LLMs)中不想要的知识。本文从优化角度审视机器遗忘,将其建模为一个带正则化的多任务优化问题,其中一项任务优化遗忘目标,另一项任务保持模型性能。我们提出一种归一化梯度差(NGDiff)算法,实现对两类目标之间权衡的更好控制,并集成一种新的自动学习率调度器。我们提供了理论分析,并在TOFU和MUSE数据集上实证证明了NGDiff在当前主流遗忘方法中的卓越表现,且训练过程稳定。

原文摘要 · Abstract (English)

Machine unlearning has been used to remove unwanted knowledge acquired by large language models (LLMs). In this paper, we examine machine unlearning from an optimization perspective, framing it as a regularized multi-task optimization problem, where one task optimizes a forgetting objective and another optimizes the model performance. In particular, we introduce a normalized gradient difference (NGDiff) algorithm, enabling us to have better control over the trade-off between the objectives, while integrating a new, automatic learning rate scheduler. We provide a theoretical analysis and empirically demonstrate the superior performance of NGDiff among state-of-the-art unlearning methods on the TOFU and MUSE datasets while exhibiting stable training.

机器遗忘多任务优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。