现有机器遗忘评估不靠谱,真实场景下模型可能没真正忘记数据。
Are We Truly Forgetting? A Critical Re-examination of Machine Unlearning Evaluation Protocols
- 用大尺度表征评估替代旧有的小规模日志评分,更真实反映遗忘效果。
- 顶尖遗忘方法要么严重降低模型表征质量,要么只改分类器却保留原始特征。
- 新设计语义相似类别的遗忘测试,逼模型真正改变特征表示,适合真实隐私场景。
机器遗忘旨在移除训练模型中的特定数据点,同时保持对保留数据的性能,以满足隐私或法律要求。尽管重要,现有评估多聚焦于小规模场景下的日志级指标,可能导致对遗忘方法在真实场景中安全性的误判。本文开展全面评估,在大规模场景下采用基于表征的评价方式,检验遗忘方法是否真正从模型表征层面消除目标数据。分析显示,当前最先进的遗忘方法要么显著劣化模型表征质量,要么仅修改分类器,从而在日志指标上表现优异,却仍保留与原模型相似的特征表示。此外,我们引入一种新评估场景:遗忘类别与下游任务类别具有语义相似性,要求特征表示与原模型显著分离,实现更严格的表征层面评估。希望本基准能成为真实条件下评估遗忘算法的标准协议。
原文摘要 · Abstract (English)
Machine unlearning is a process to remove specific data points from a trained model while maintaining the performance on the retain data, addressing privacy or legal requirements. Despite its importance, existing unlearning evaluations tend to focus on logit-based metrics under small-scale scenarios. We observe that this could lead to a false sense of security in unlearning approaches under real-world scenarios. In this paper, we conduct a comprehensive evaluation that employs representation-based evaluations of the unlearned model under large-scale scenarios to verify whether the unlearning approaches truly eliminate the targeted data from the model's representation perspective. Our analysis reveals that current state-of-the-art unlearning approaches either completely degrade the representational quality of the unlearned model or merely modify the classifier, thereby achieving superior logit-based performance while maintaining representational similarity to the original model. Furthermore, we introduce a novel unlearning evaluation scenario in which the forgetting classes exhibit semantic similarity to downstream task classes, necessitating that feature representations diverge significantly from those of the original model, thus enabling a more thorough evaluation from a representation perspective. We hope our benchmark will serve as a standardized protocol for evaluating unlearning algorithms under realistic conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。