arXiv:2409.13864cs.LGcs.CR2024-09被引 18

提出两种新型持续学习后门攻击,可长期潜伏并绕过防御。

Persistent Backdoor Attacks in Continual Learning

  • 设计两种低暴露攻击:盲任务与潜在任务后门,仅微调损失或单任务训练。
  • 在多种触发器下保持高成功率,跨不同持续学习算法稳定有效。
  • 能成功规避SentiNet、I-BAU等主流防御机制,适合安全研究者关注。

后门攻击对神经网络构成重大威胁,使攻击者能在特定输入上操控模型输出,尤其在关键应用中后果严重。尽管后门攻击已在多种场景被研究,但其在持续学习中的实际可行性与持久性仍缺乏关注,尤其是新数据分布不断学习和融合时,模型参数的持续更新如何影响攻击效果尚不明确。为填补这一空白,我们提出两种持久性后门攻击——盲任务后门与潜在任务后门,均仅需极小的恶意影响。盲任务后门通过隐式修改损失计算实现,无需直接控制训练过程;潜在任务后门仅干扰单一任务训练,其余任务均正常进行。我们在多种配置下评估了这些攻击,涵盖静态、动态、物理及语义触发器。结果表明,两种攻击在不同持续学习算法中均保持高成功率,且能有效规避SentiNet与I-BAU等先进防御机制。

原文摘要 · Abstract (English)

Backdoor attacks pose a significant threat to neural networks, enabling adversaries to manipulate model outputs on specific inputs, often with devastating consequences, especially in critical applications. While backdoor attacks have been studied in various contexts, little attention has been given to their practicality and persistence in continual learning, particularly in understanding how the continual updates to model parameters, as new data distributions are learned and integrated, impact the effectiveness of these attacks over time. To address this gap, we introduce two persistent backdoor attacks-Blind Task Backdoor and Latent Task Backdoor-each leveraging minimal adversarial influence. Our blind task backdoor subtly alters the loss computation without direct control over the training process, while the latent task backdoor influences only a single task's training, with all other tasks trained benignly. We evaluate these attacks under various configurations, demonstrating their efficacy with static, dynamic, physical, and semantic triggers. Our results show that both attacks consistently achieve high success rates across different continual learning algorithms, while effectively evading state-of-the-art defenses, such as SentiNet and I-BAU.

后门攻击持续学习模型安全防御绕过

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。