arXiv:2501.19403cs.LGcs.AI2025-01被引 2

用不确定性量化发现模型‘假遗忘’问题,提出新评估方法与框架。

Tackling Fake Forgetting through Uncertainty Quantification

  • 引入符合预测机制,识别被误判为已遗忘的数据点。
  • 新指标CR能更可靠评估遗忘质量,实验显示效果更优。
  • 适合关注模型可解释性与安全性的研究者使用。

机器遗忘旨在消除特定数据对训练模型的影响。尽管遗忘准确率是广泛使用的评估指标,但其无法衡量遗忘的可靠性。本文发现,即使某些数据点在遗忘准确率下被错误分类,它们的真实标签仍可能保留在基于不确定性量化的符合预测集合中,这种现象被称为‘假遗忘’。为此,我们提出一种受符合预测启发的新指标CR,可更可靠地评估遗忘质量。基于此,我们进一步提出一个名为CPU的遗忘框架,将符合预测融入Carlini & Wagner对抗攻击损失中,使真实标签从符合预测集合中被有效移除。在图像分类任务上的大量实验表明,所提指标有效,且该框架实现了更优的遗忘质量。代码已公开于https://github.com/TIML-Group/Conformal-Prediction-Unlearning。

原文摘要 · Abstract (English)

Machine unlearning seeks to remove the influence of specified data from a trained model. While the unlearning accuracy provides a widely used metric for assessing unlearning performance, it falls short in assessing the reliability of forgetting. In this paper, we find that the forgetting data points misclassified by unlearning accuracy still have their ground truth labels included in the conformal prediction set from the uncertainty quantification perspective, leading to a phenomenon we term fake forgetting. To address this issue, we propose a novel metric CR, inspired by conformal prediction, that offers a more reliable assessment of forgetting quality. Building on these insights, we further propose an unlearning framework CPU that incorporates conformal prediction into the Carlini & Wagner adversarial attack loss, enabling the ground truth label to be effectively removed from the conformal prediction set. Through extensive experiments on image classification tasks, we demonstrate both the effectiveness of our proposed metric and the superior forgetting quality achieved by our framework. Code is available at https://github.com/TIML-Group/Conformal-Prediction-Unlearning.

机器遗忘不确定性量化符合预测安全模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。