arXiv:2510.13322cs.CRcs.AI2025-10AAAI

提出可主动清除的后门攻击,攻完即删不留痕迹。

Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning

  • 设计双层优化触发器,兼顾攻击成功率与易清除性。
  • 在CIFAR-10和ImageNet上保持高攻击成功率,且后门可彻底移除。
  • 适合研究模型可撤销攻击与防御机制的安全团队。

后门攻击因隐蔽性强、持久存在而持续威胁深度神经网络。尽管已有研究尝试利用模型遗忘机制增强后门隐藏,但现有攻击仍留下可被静态分析检测的持久痕迹。本文首次提出可撤销后门攻击范式,使后门在达成攻击目标后可主动、彻底地被移除。我们将可撤销后门攻击中的触发器优化建模为双层优化问题:通过模拟注入与遗忘过程,优化触发器生成器以实现高攻击成功率(ASR),同时确保后门可通过遗忘机制轻松消除。为缓解注入与移除目标间的优化冲突,我们采用确定性划分中毒样本与遗忘样本,减少采样方差,并进一步应用投影冲突梯度(PCGrad)技术解决剩余梯度冲突。在CIFAR-10和ImageNet上的实验表明,该方法在攻击成功率方面达到当前最优水平,同时支持有效移除后门行为。本工作开辟了后门攻击研究的新方向,也为机器学习系统安全带来新挑战。

原文摘要 · Abstract (English)

Backdoor attacks pose a persistent security risk to deep neural networks (DNNs) due to their stealth and durability. While recent research has explored leveraging model unlearning mechanisms to enhance backdoor concealment, existing attack strategies still leave persistent traces that may be detected through static analysis. In this work, we introduce the first paradigm of revocable backdoor attacks, where the backdoor can be proactively and thoroughly removed after the attack objective is achieved. We formulate the trigger optimization in revocable backdoor attacks as a bilevel optimization problem: by simulating both backdoor injection and unlearning processes, the trigger generator is optimized to achieve a high attack success rate (ASR) while ensuring that the backdoor can be easily erased through unlearning. To mitigate the optimization conflict between injection and removal objectives, we employ a deterministic partition of poisoning and unlearning samples to reduce sampling-induced variance, and further apply the Projected Conflicting Gradient (PCGrad) technique to resolve the remaining gradient conflicts. Experiments on CIFAR-10 and ImageNet demonstrate that our method maintains ASR comparable to state-of-the-art backdoor attacks, while enabling effective removal of backdoor behavior after unlearning. This work opens a new direction for backdoor attack research and presents new challenges for the security of machine learning systems.

后门攻击模型遗忘可撤销安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。