模型遗忘后仍可能识别被删数据的微扰版本,存在隐私泄露风险。
The Unseen Threat: Residual Knowledge in Machine Unlearning under Perturbed Samples
- 提出残余知识概念,揭示遗忘后模型对微扰数据的异常识别能力
- 实验显示现有方法在视觉任务中普遍存在残余知识,且无法完全消除
- 设计RURK方法通过微调惩罚模型重识别能力,有效缓解该风险
机器遗忘为避免全量重训练提供了实用方案,通过近似消除特定用户数据的影响。现有方法通过统计不可区分性验证遗忘效果,但此类保证在输入遭对抗扰动时不再成立。尤其当被遗忘样本受到轻微扰动后,已遗忘模型仍可正确识别——而重新训练模型反而无法识别,暴露出新隐私风险:被遗忘数据的信息可能残留在其局部邻域。本文将此现象形式化为残余知识,并证明其在高维设置下不可避免。为此,提出名为RURK的微调策略,通过惩罚模型对扰动后的遗忘样本的重识别能力来缓解该风险。在深度神经网络和视觉基准上的实验表明,残余知识广泛存在于现有遗忘方法中,而本方法能有效防止其发生。
原文摘要 · Abstract (English)
Machine unlearning offers a practical alternative to avoid full model re-training by approximately removing the influence of specific user data. While existing methods certify unlearning via statistical indistinguishability from re-trained models, these guarantees do not naturally extend to model outputs when inputs are adversarially perturbed. In particular, slight perturbations of forget samples may still be correctly recognized by the unlearned model - even when a re-trained model fails to do so - revealing a novel privacy risk: information about the forget samples may persist in their local neighborhood. In this work, we formalize this vulnerability as residual knowledge and show that it is inevitable in high-dimensional settings. To mitigate this risk, we propose a fine-tuning strategy, named RURK, that penalizes the model's ability to re-recognize perturbed forget samples. Experiments on vision benchmarks with deep neural networks demonstrate that residual knowledge is prevalent across existing unlearning methods and that our approach effectively prevents residual knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。