提出仅需模型输出标签的隐私攻击,可高效检测数据是否被机器遗忘。
Apollo: A Posteriori Label-Only Membership Inference Attack Towards Machine Unlearning
- 攻击者仅需未学习模型的标签输出,无需访问原始模型。
- 在严格威胁模型下仍能高精度判断样本是否被删除。
- 揭示机器遗忘技术潜在隐私风险,适合关注模型安全的研究者。
机器遗忘(MU)旨在高效更新机器学习模型,以响应删除训练样本及其影响的请求,而无需从头重新训练。尽管MU可用于保护隐私并满足合规要求,但它也可能增加模型的攻击面。现有针对MU的隐私推断攻击通常假设攻击者同时拥有未学习模型和原始模型,这限制了其在真实场景中的可行性。本文提出一种新型隐私攻击——后验仅标签成员推理攻击(Apollo),该攻击在严格威胁模型下仅需访问未学习模型的标签输出,即可推断某数据样本是否已被遗忘。我们证明,相较于先前攻击,本方法所需模型访问更少,但仍能对被遗忘样本的成员身份实现较高精度的推断。
原文摘要 · Abstract (English)
Machine Unlearning (MU) aims to update Machine Learning (ML) models following requests to remove training samples and their influences on a trained model efficiently without retraining the original ML model from scratch. While MU itself has been employed to provide privacy protection and regulatory compliance, it can also increase the attack surface of the model. Existing privacy inference attacks towards MU that aim to infer properties of the unlearned set rely on the weaker threat model that assumes the attacker has access to both the unlearned model and the original model, limiting their feasibility toward real-life scenarios. We propose a novel privacy attack, A Posteriori Label-Only Membership Inference Attack towards MU, Apollo, that infers whether a data sample has been unlearned, following a strict threat model where an adversary has access to the label-output of the unlearned model only. We demonstrate that our proposed attack, while requiring less access to the target model compared to previous attacks, can achieve relatively high precision on the membership status of the unlearned samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。