arXiv:2508.07458cs.LG2025-08被引 3

首次揭示模型遗忘后预测不确定性漏洞,可被恶意攻击操纵

Towards Unveiling Predictive Uncertainty Vulnerabilities in the Context of the Right to Be Forgotten

  • 提出针对预测不确定性的新型恶意遗忘攻击方法
  • 实验表明该攻击比传统误分类攻击更有效,且现有防御无效
  • 适合关注隐私安全与模型可信性的研究人员

当前已有多种不确定性量化方法为深度学习模型的预测结果提供置信度和概率估计。与此同时,随着对删除权需求的增长,机器遗忘技术被广泛研究,可在不从头训练的情况下移除敏感数据对预训练模型的影响。然而,此类生成的预测不确定性在面对专门的恶意遗忘攻击时是否存在漏洞,尚未被探索。为此,我们首次提出一类针对预测不确定性的恶意遗忘攻击,攻击者旨在操控特定预测结果的不确定性。我们设计了新的优化框架并进行了大量实验,涵盖黑盒场景。值得注意的是,实验结果显示,我们的攻击在操纵预测不确定性方面比专注于标签误分类的传统攻击更具效果,且现有的传统攻击防御手段对本攻击无效。

原文摘要 · Abstract (English)

Currently, various uncertainty quantification methods have been proposed to provide certainty and probability estimates for deep learning models' label predictions. Meanwhile, with the growing demand for the right to be forgotten, machine unlearning has been extensively studied as a means to remove the impact of requested sensitive data from a pre-trained model without retraining the model from scratch. However, the vulnerabilities of such generated predictive uncertainties with regard to dedicated malicious unlearning attacks remain unexplored. To bridge this gap, for the first time, we propose a new class of malicious unlearning attacks against predictive uncertainties, where the adversary aims to cause the desired manipulations of specific predictive uncertainty results. We also design novel optimization frameworks for our attacks and conduct extensive experiments, including black-box scenarios. Notably, our extensive experiments show that our attacks are more effective in manipulating predictive uncertainties than traditional attacks that focus on label misclassifications, and existing defenses against conventional attacks are ineffective against our attacks.

模型安全机器遗忘不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。