arXiv:2504.05058cs.CL2025-04中稿 · COLM被引 12

不同数据点的遗忘难度差异巨大,高频知识更难清除。

Not All Data Are Unlearned Equally

  • 按知识出现频率分层处理,而非统一对待。
  • 高频知识在预训练中出现越多,越难从模型中删除。
  • 大模型下传统评估方法会严重低估遗忘效果。

机器遗忘旨在从已训练模型中移除特定数据点所包含的知识。在大型语言模型(LLM)背景下,这一问题因隐私需求而受到关注,例如移除关于命名实体的知识。现有方法普遍假设所有待遗忘数据具有相同难度,即移除‘蒙特利尔是加拿大城市’与移除本文第一作者电话号码同等困难。本文揭示该假设不成立:我们发现遗忘成功率高度依赖于目标知识在预训练数据中的出现频率——频率越高,越难遗忘。此外,我们发现概率评估与生成评估之间存在偏差,且该偏差随模型规模增大而加剧。实验表明,亟需考虑模型训练数据分布的新方法与更合理的评估体系。

原文摘要 · Abstract (English)

Machine unlearning is concerned with the task of removing knowledge learned from particular data points from a trained model. In the context of large language models (LLMs), unlearning has recently received increased attention, particularly for removing knowledge about named entities from models for privacy purposes. While various approaches have been proposed to address the unlearning problem, most existing approaches treat all data points to be unlearned equally, i.e., unlearning that Montreal is a city in Canada is treated exactly the same as unlearning the phone number of the first author of this paper. In this work, we show that this all data is equal assumption does not hold for LLM unlearning. We study how the success of unlearning depends on the frequency of the knowledge we want to unlearn in the pre-training data of a model and find that frequency strongly affects unlearning, i.e., more frequent knowledge is harder to unlearn. Additionally, we uncover a misalignment between probability and generation-based evaluations of unlearning and show that this problem worsens as models become larger. Overall, our experiments highlight the need for better evaluation practices and novel methods for LLM unlearning that take the training data of models into account.

机器遗忘大模型隐私保护评估偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。