arXiv:2506.19486cs.LGcs.AI2025-06IJCAI被引 2

未学习模型可泄露遗忘数据的类别信息,攻击者无需原始模型即可还原隐私。

Recalling The Forgotten Class Memberships: Unlearned Models Can Be Noisy Labelers to Leak Privacy

  • 用教师-学生架构让未学习模型充当噪声标签源,实现知识迁移。
  • 在多个真实数据集上成功恢复遗忘样本的类别归属,准确率超基准方法。
  • 适用于研究模型遗忘安全性的研究人员,揭示当前技术的隐私漏洞。

机器遗忘(MU)技术可按请求移除特定数据对训练模型的影响。尽管MU技术发展迅速,其潜在漏洞仍待深入探索,可能引发本应被遗忘的信息泄露风险。现有少数关于MU攻击的研究需访问包含隐私数据的原始模型,违背了MU的核心隐私保护目标。为此,我们首次提出无需原始模型即可从已遗忘模型(ULMs)中重构被遗忘实例类别信息的新方法。具体地,设计了一种基于教师-学生知识蒸馏框架的成员回忆攻击(MRA),使ULMs作为噪声标签生成器,将知识传递给学生模型,进而转化为带噪声标签的学习问题以推断遗忘实例的真实标签。在多种主流MU方法与真实数据集上的大量实验表明,所提MRA策略能高效恢复遗忘实例的类别归属。本研究为未来MU安全性评估建立了基准。

原文摘要 · Abstract (English)

Machine Unlearning (MU) technology facilitates the removal of the influence of specific data instances from trained models on request. Despite rapid advancements in MU technology, its vulnerabilities are still underexplored, posing potential risks of privacy breaches through leaks of ostensibly unlearned information. Current limited research on MU attacks requires access to original models containing privacy data, which violates the critical privacy-preserving objective of MU. To address this gap, we initiate an innovative study on recalling the forgotten class memberships from unlearned models (ULMs) without requiring access to the original one. Specifically, we implement a Membership Recall Attack (MRA) framework with a teacher-student knowledge distillation architecture, where ULMs serve as noisy labelers to transfer knowledge to student models. Then, it is translated into a Learning with Noisy Labels (LNL) problem for inferring the correct labels of the forgetting instances. Extensive experiments on state-of-the-art MU methods with multiple real datasets demonstrate that the proposed MRA strategy exhibits high efficacy in recovering class memberships of unlearned instances. As a result, our study and evaluation have established a benchmark for future research on MU vulnerabilities.

机器遗忘隐私泄露知识蒸馏攻击分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。