arXiv:2512.00272cs.LGcs.AI2025-12中稿 · publication at the…被引 1

提出WARP防御机制,让模型遗忘数据时更安全。

WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols

  • 利用神经网络对称性重参数化,隐藏遗忘数据痕迹。
  • 黑盒攻击下隐私优势提升64%,白盒达92%。
  • 兼容多种算法,适合注重数据隐私的研究者。

近似机器遗忘旨在高效移除特定数据对训练模型的影响,是全量重训练的实用替代方案。然而,该方法引入隐私风险:攻击者通过比较遗忘前后的模型,可实施成员推断或数据重建。我们发现这些漏洞源于两方面:遗忘样本梯度范数过大,以及遗忘后参数与原模型过于接近。为此,我们设计了针对遗忘场景的成员推断和重建攻击,验证了多项前沿方法(如NGP、SCRUB)仍易受攻击。为缓解泄漏,提出WARP——一种即插即用的传输防御机制,利用神经网络对称性降低遗忘样本梯度能量,增加参数分散度,同时保持预测性能。该重参数化过程混淆遗忘数据信号,使攻击者更难区分遗忘样本与非成员,或通过重建恢复数据。在六种遗忘算法上,该方法均实现稳定隐私提升,黑盒下对抗优势(AUC)下降最高达64%,白盒下达92%,且保留数据准确率不受影响。结果表明,传输机制是降低近似遗忘中攻击成功率的通用有效工具。

原文摘要 · Abstract (English)

Approximate machine unlearning aims to efficiently remove the influence of specific data points from a trained model, offering a practical alternative to full retraining. However, it introduces privacy risks: an adversary with access to pre- and post-unlearning models can exploit their differences for membership inference or data reconstruction. We show these vulnerabilities arise from two factors: large gradient norms of forget-set samples and the close proximity of unlearned parameters to the original model. To demonstrate their severity, we propose unlearning-specific membership inference and reconstruction attacks, showing that several state-of-the-art methods (e.g., NGP, SCRUB) remain vulnerable. To mitigate this leakage, we introduce WARP, a plug-and-play teleportation defense that leverages neural network symmetries to reduce forget-set gradient energy and increase parameter dispersion while preserving predictions. This reparameterization obfuscates the signal of forgotten data, making it harder for attackers to distinguish forgotten samples from non-members or recover them via reconstruction. Across six unlearning algorithms, our approach achieves consistent privacy gains, reducing adversarial advantage (AUC) by up to 64% in black-box and 92% in white-box settings, while maintaining accuracy on retained data. These results highlight teleportation as a general tool for reducing attack success in approximate unlearning.

模型遗忘隐私保护对抗攻击重参数化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。