arXiv:2601.20325cs.CRcs.CV2026-01中稿 · ICASSP 2026被引 1

提出首个防御反向重构隐私数据的遗忘防护方案

UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion

  • 在余弦空间引入方向性扰动并约束调节
  • 实现遗忘效果、精度与隐私保护的平衡
  • 适合关注模型隐私安全的研究者

机器遗忘是一种新兴技术,旨在消除特定数据对训练模型的影响,从而增强隐私保护。然而,近期研究发现关键隐私漏洞:攻击者可利用遗忘反演重构本应被删除的数据。尽管威胁严重,现有防御手段仍匮乏。为此,我们提出UnlearnShield,首个专为抵御遗忘反演设计的防御机制。该方法在余弦表示空间引入方向性扰动,并通过约束模块协同调节,以同时保障模型精度与遗忘有效性,从而降低反演风险并维持模型可用性。实验表明,该方案在隐私保护、准确率与遗忘性能间实现了良好权衡。

原文摘要 · Abstract (English)

Machine unlearning is an emerging technique that aims to remove the influence of specific data from trained models, thereby enhancing privacy protection. However, recent research has uncovered critical privacy vulnerabilities, showing that adversaries can exploit unlearning inversion to reconstruct data that was intended to be erased. Despite the severity of this threat, dedicated defenses remain lacking. To address this gap, we propose UnlearnShield, the first defense specifically tailored to counter unlearning inversion. UnlearnShield introduces directional perturbations in the cosine representation space and regulates them through a constraint module to jointly preserve model accuracy and forgetting efficacy, thereby reducing inversion risk while maintaining utility. Experiments demonstrate that it achieves a good trade-off among privacy protection, accuracy, and forgetting.

隐私保护模型遗忘反演攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。