arXiv:2506.00688cs.LGcs.AI2025-06被引 7

现有大模型遗忘评估方法不可靠,易误判遗忘效果。

Existing Large Language Model Unlearning Evaluations Are Inconclusive

  • 发现评估时注入新信息,干扰真实遗忘效果判断。
  • 不同任务结果差异大,现有评估缺乏通用性。
  • 多依赖虚假关联,结果难解释,适合研究评估规范者看。

机器遗忘旨在从大语言模型中移除敏感或不想要的数据。然而,近期研究指出遗忘往往流于表面,被删除的知识容易被恢复。本文批判性审视了标准遗忘评估方法,揭示其关键缺陷:首先,部分评估在测试阶段向模型引入大量新信息,可能掩盖真实的遗忘表现;其次,评估结果在不同任务间差异显著,削弱了现有评估流程的普适性;最后,许多评估依赖虚假相关性,导致结果难以信任与解释。综合来看,当前评估协议可能同时高估和低估遗忘成效。为此,我们提出两个未来评估原则:最小化信息注入与下游任务意识。通过一系列针对性实验验证,表明违背任一原则均会导致误导性结论。

原文摘要 · Abstract (English)

Machine unlearning aims to remove sensitive or undesired data from large language models. However, recent studies suggest that unlearning is often shallow, claiming that removed knowledge can easily be recovered. In this work, we critically examine standard unlearning evaluation practices and uncover key limitations that shake our trust in those findings. First, we show that some evaluations introduce substantial new information into the model, potentially masking true unlearning performance by re-teaching the model during testing. Second, we demonstrate that evaluation outcomes vary significantly across tasks, undermining the generalizability of current evaluation routines. Finally, we find that many evaluations rely on spurious correlations, making their results difficult to trust and interpret. Taken together, these issues suggest that current evaluation protocols may both overstate and understate unlearning success. To address this, we propose two principles for future unlearning evaluations: minimal information injection and downstream task awareness. We validate these principles through a series of targeted experiments, showing how violations of each can lead to misleading conclusions.

大模型遗忘评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。