提出新评估框架,揭示现有机器遗忘方法存在可被区分的漏洞。
Mirror Mirror on the Wall, Have I Forgotten it All? A New Framework for Evaluating Machine Unlearning
- 用对抗性区分算法检验遗忘模型与重训练对照模型差异
- 证明现有方法无法满足计算意义上的完全遗忘
- 指出差分隐私方法虽能实现但严重损害模型性能
机器遗忘旨在让模型表现得像从未学习过某些数据。我们实证发现,攻击者能通过文献中的成员推断分数和KL散度等指标,有效区分遗忘模型与重训练对照模型。为此提出新的形式化定义——计算遗忘:若攻击者无法以高于随机猜测的准确率区分两者,则视为达成计算遗忘。该定义表明,对于熵学习算法,不存在确定性计算遗忘方法。同时发现,基于差分隐私的遗忘方法虽能满足计算遗忘,但会导致模型性能大幅下降。这些结果表明当前方法在理论上均未真正实现完全遗忘,并指出了未来研究方向。
原文摘要 · Abstract (English)
Machine unlearning methods take a model trained on a dataset and a forget set, then attempt to produce a model as if it had only been trained on the examples not in the forget set. We empirically show that an adversary is able to distinguish between a mirror model (a control model produced by retraining without the data to forget) and a model produced by an unlearning method across representative unlearning methods from the literature. We build distinguishing algorithms based on evaluation scores in the literature (i.e. membership inference scores) and Kullback-Leibler divergence. We propose a strong formal definition for machine unlearning called computational unlearning. Computational unlearning is defined as the inability for an adversary to distinguish between a mirror model and a model produced by an unlearning method. If the adversary cannot guess better than random (except with negligible probability), then we say that an unlearning method achieves computational unlearning. Our computational unlearning definition provides theoretical structure to prove unlearning feasibility results. For example, our computational unlearning definition immediately implies that there are no deterministic computational unlearning methods for entropic learning algorithms. We also explore the relationship between differential privacy (DP)-based unlearning methods and computational unlearning, showing that DP-based approaches can satisfy computational unlearning at the cost of an extreme utility collapse. These results demonstrate that current methodology in the literature fundamentally falls short of achieving computational unlearning. We conclude by identifying several open questions for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。