arXiv:2505.16831cs.CLcs.AI2025-05中稿 · appear被引 37

发现大模型删数据后信息可轻易恢复,揭示现有评估方法的漏洞。

Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs

  • 从表征层面分析模型遗忘效果,用多种指标检测信息残留。
  • 发现多数删数据操作只是压制信息,微调后即可快速恢复原行为。
  • 提出新评估框架,适合研究模型隐私安全与可信删除的学者。

大型语言模型的去学习旨在移除特定数据,但其有效性通常通过任务级指标(如准确率和困惑度)评估。我们发现这些指标具有误导性:模型看似已遗忘,但通过极少量微调即可轻松恢复原有行为,表明信息仅被抑制而非真正擦除。为此,我们提出一种表征层面分析框架,包含主成分分析相似性与偏移、中心核对齐(CKA)及费舍尔信息,并以平均主成分距离作为综合度量,用于衡量表征漂移。在多种去学习方法、数据领域和大模型上应用该框架,识别出四种基于可逆性和灾难性程度的遗忘模式。我们对比了恢复策略,发现重学效率取决于数据来源。还发现不可逆且非灾难性的遗忘极为困难。通过探测去学习的极限,我们识别出一个看似不可逆的针对性遗忘案例,为更鲁棒的擦除算法提供洞见。总体而言,研究揭示了当前评估体系的缺失,并建立了可信去学习的表征基础。

原文摘要 · Abstract (English)

Unlearning in large language models (LLMs) aims to remove specified data, but its efficacy is typically assessed with task-level metrics like accuracy and perplexity. We show that these metrics can be misleading, as models can appear to forget while their original behavior is easily restored through minimal fine-tuning. This \emph{reversibility} suggests that information is merely suppressed, not genuinely erased. To address this critical evaluation gap, we introduce a \emph{representation-level analysis framework}. Our toolkit comprises PCA similarity and shift, centered kernel alignment (CKA), and Fisher information, complemented by a summary metric, the mean PCA distance, to measure representational drift. Applying this framework across multiple unlearning methods, data domains, and LLMs, we identify four distinct forgetting regimes based on their \emph{reversibility} and \emph{catastrophicity}. We compare recovery strategies and show that relearning efficiency relies on the data source. We also find that irreversible, non-catastrophic forgetting is exceptionally challenging. By probing unlearning limits, we identify a case of seemingly irreversible, targeted forgetting, offering insights for more robust erasure algorithms. Overall, our findings expose a gap in current evaluation and establish a representation-level foundation for trustworthy unlearning.

大模型去学习可逆性表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。