arXiv:2507.10886cs.LGcs.AI2025-07

提出防御恶意删除请求的模型保护方法,防止性能下降。

How to Protect Models against Adversarial Unlearning?

  • 设计新机制抵御恶意发起的删除请求
  • 在多种删除策略下保持模型性能稳定
  • 适合需合规删除数据的AI系统使用

AI模型需支持删除功能以满足《人工智能法案》或GDPR等法律要求,也用于清除有毒内容、去偏、应对恶意样本或数据分布变化。然而,知识移除可能导致模型性能下降。本文研究对抗性删减问题:恶意方故意发送删除请求以最大程度损害模型性能。我们发现该现象及其攻击能力受模型架构及待删数据选择策略影响显著。本文提出一种新方法,可同时抵御自发性删除与恶意攻击带来的性能衰退。

原文摘要 · Abstract (English)

AI models need to be unlearned to fulfill the requirements of legal acts such as the AI Act or GDPR, and also because of the need to remove toxic content, debiasing, the impact of malicious instances, or changes in the data distribution structure in which a model works. Unfortunately, removing knowledge may cause undesirable side effects, such as a deterioration in model performance. In this paper, we investigate the problem of adversarial unlearning, where a malicious party intentionally sends unlearn requests to deteriorate the model's performance maximally. We show that this phenomenon and the adversary's capabilities depend on many factors, primarily on the backbone model itself and strategy/limitations in selecting data to be unlearned. The main result of this work is a new method of protecting model performance from these side effects, both in the case of unlearned behavior resulting from spontaneous processes and adversary actions.

模型安全隐私保护对抗性攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。