arXiv:2503.20925cs.CVcs.AI2025-03被引 1

提出一种能防御多种触发器的后置鲁棒检测方法

Prototype Guided Backdoor Defense

  • 利用激活空间几何位移设计新型净化损失
  • 对各类触发器(含语义触发)均实现更高防御效果
  • 特别适用于生成式攻击下的名人图像防御

深度学习模型易受后门攻击,攻击者通过在少量训练数据中加入触发器导致误分类。现有触发器包括无需修改图像即可实现的语义触发,而生成式AI使伪造样本更易生成。跨类型触发的鲁棒性对有效防御至关重要。本文提出原型引导的后门防御(PGBD),一种可扩展至多种触发类型的鲁棒后置防御方法。PGBD利用激活空间中的几何位移来惩罚向触发器方向的移动,通过后置微调阶段引入新颖的净化损失实现。该几何方法可轻松适配各类攻击。PGBD在所有测试场景下均表现更优,并首次实现了对名人面部图像上新语义攻击的有效防御。

原文摘要 · Abstract (English)

Deep learning models are susceptible to {\em backdoor attacks} involving malicious attackers perturbing a small subset of training data with a {\em trigger} to causes misclassifications. Various triggers have been used, including semantic triggers that are easily realizable without requiring the attacker to manipulate the image. The emergence of generative AI has eased the generation of varied poisoned samples. Robustness across types of triggers is crucial to effective defense. We propose Prototype Guided Backdoor Defense (PGBD), a robust post-hoc defense that scales across different trigger types, including previously unsolved semantic triggers. PGBD exploits displacements in the geometric spaces of activations to penalize movements toward the trigger. This is done using a novel sanitization loss of a post-hoc fine-tuning step. The geometric approach scales easily to all types of attacks. PGBD achieves better performance across all settings. We also present the first defense against a new semantic attack on celebrity face images. Project page: \hyperlink{https://venkatadithya9.github.io/pgbd.github.io/}{this https URL}.

后门攻击防御机制生成式攻击几何方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。