从持续学习视角重新定义后门消除,实现彻底清除攻击痕迹。
Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in Continual Learning

- 将后门学习与消除视为连续三阶段过程,利用灾难性遗忘机制设计条件。
- 提出BI-BAU方法,通过盲反演生成满足消除条件的对抗样本。
- 可应对未知目标类和多模态任务,适合真实场景中受污染模型修复。
现有研究发现,当前后门防御方法鲁棒性有限,常无法抵御特定攻击。更令人担忧的是,主流安全调优策略仅提供表面保护,未能完全消除后门影响。本文从持续学习视角,将后门学习与消除建模为一个序列化的三阶段过程。在此框架下,我们正式定义了完整的后门消除,并基于灾难性遗忘机制推导出实现它的必要条件。据此,我们提出盲反演-后门对抗去学习(BI-BAU),将满足消除条件的对抗样本生成问题转化为盲反演问题。通过将对抗训练的双层优化嵌入期望最大化(EM)算法框架,以优化最大后验概率(MAP)目标求解。此外,该方法被扩展至未知目标类的无目标对抗场景及多模态对比学习任务,增强了其在实际部署中的适用性。大量实验表明,本方法对多种后门攻击具有广泛适应性,能有效且彻底地从后门模型中消除后门效应。
原文摘要 · Abstract (English)
Existing studies reveal that current backdoor defenses exhibit limited robustness and often fail against specific types of attacks. More concerningly, prevailing safety tuning strategies tend to provide only superficial safety protection, as they fall short of completely eliminating the backdoor effects. In this work, we present a novel formulation of backdoor learning and unlearning as a sequential, three-stage process from a continual learning perspective. Within this framework, we formally define complete backdoor unlearning and further derive the necessary conditions for achieving it based on the mechanism of catastrophic forgetting. Guided by these insights, we propose Blind Inversion-Backdoor Adversarial Unlearning (BI-BAU), which formulates the generation of adversarial examples satisfying the unlearning conditions as a blind inversion problem. We solve this by integrating the bi-level optimization process of adversarial training into an Expectation-Maximization (EM) algorithm framework to optimize the maximum a posteriori (MAP) objective. Furthermore, BI-BAU is extended to untargeted adversarial scenarios with unknown target classes, as well as to multi-modal contrastive learning tasks, enhancing its applicability to real-world deployment scenarios where pre-trained models may be compromised. Extensive experiments demonstrate that our method exhibits general applicability across a wide spectrum of backdoor attacks and can effectively and thoroughly eliminate the backdoor effects from a backdoor model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。