arXiv:2606.16110cs.LG2026-06

提出首个可实用的模型遗忘审计框架,验证算法是否真删除数据影响。

Auditing Machine Unlearning: A Systematic Research on Whether Models Truly Forget

论文配图:Auditing Machine Unlearning: A Systematic Research on Whether Models Truly Forget
图 1 · 摘自论文原文
  • 基于'无知证明'思想,无需重训或大量影子模型即可审计。
  • 在6个数据集、10种方法上验证:微调类方法有效,优化反向类失败。
  • 可识别虚假遗忘,适用于大语言模型,具强泛化性。

机器遗忘因隐私关切与法规要求受到广泛关注,但如何审计算法是否真正抹除特定数据的影响仍属开放挑战。本文首次系统研究现有遗忘算法能否真正遗忘指定数据,提出首个实用且通用的审计框架,灵感来自‘无知证明’概念。该框架克服了现有方法的实用性缺陷:无需从头重训基线,避免训练大量影子模型,也无需对原始训练过程进行侵入式干预。为验证框架有效性,我们首先开展验证实验,确认其正确性和完备性;随后在六个数据集和十种代表性遗忘方法上进行综合实验。结果表明,该框架能可靠区分成功与失败的遗忘。特别地,重训与微调类方法即使目标数据仍在原数据集中,仍可实现有效遗忘;而去优化类方法未能实现真正遗忘,反而降低模型性能;基于Fisher/Hessian的方法即使提供形式化认证也未能成功遗忘。此外,框架对虚假遗忘尝试具有鲁棒性,并能良好泛化至大语言模型。

原文摘要 · Abstract (English)

Machine unlearning has been extensively studied in response to growing privacy concerns and regulatory requirements. However, auditing whether unlearning algorithms have truly erased the influence of specific data remains an open challenge. The lack of reliable and practical auditing mechanisms can lead to critical privacy risks, such as residual information leakage. This paper initiates a systematic investigation into whether existing unlearning algorithms can truly forget the designated data. We propose the first practical and general-purpose auditing framework for machine unlearning, inspired by the concept of proof of ignorance. Our framework addresses the key practicality limitations of existing methods by eliminating the need for retraining-from-scratch baselines, avoiding the training of large numbers of shadow models, and requiring no intrusive intervention in the original training process. To evaluate the effectiveness of our framework, we first conduct validation experiments to verify its soundness and completeness. We then perform comprehensive experiments across six datasets and ten representative unlearning methods. The results demonstrate that our framework reliably distinguishes between successful and failed unlearning. In particular, we observe that retraining-based and fine-tuning-based methods can achieve effective unlearning, even when the target data remain in the original dataset. In contrast, de-optimization-based methods fail to achieve true unlearning and instead degrade the model's performance. Fisher/Hessian-based methods also fail to unlearn requested data, even formal certification is provided. Moreover, we show that our framework is robust against fake unlearning attempts and generalizes well to large language models.

机器遗忘隐私审计模型安全大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。