提出新型后门攻击Grond,让模型参数更隐蔽,逃过多重防御。
Towards Backdoor Stealthiness in Model Parameter Space
- 通过自适应模块控制参数变化,提升攻击在参数空间的隐蔽性。
- 在CIFAR-10、GTSRB等数据集上超越12种现有攻击,对抗先进防御。
- 适合关注模型安全、后门防御的开发者和研究人员参考。
当前后门攻击多聚焦于输入空间和特征空间的隐蔽性,以规避相关检测。然而,这些攻击往往仅针对特定防御机制,缺乏对多样化防御的鲁棒性。我们评估了12种主流后门攻击与17种代表性防御,发现一个关键盲点:尽管攻击在输入和特征空间中隐蔽,但在参数空间仍暴露明显特征。分析表明,这类攻击会在参数空间引入显著的后门神经元。为此,我们提出新型供应链攻击Grond,其核心模块Adversarial Backdoor Injection(ABI)能自适应地限制参数扰动,提升参数空间隐蔽性。大量实验显示,Grond在CIFAR-10、GTSRB及部分ImageNet上均优于所有12种现有攻击,包括对抗自适应防御的方法。此外,ABI可普遍增强常见攻击的有效性。
原文摘要 · Abstract (English)
Recent research on backdoor stealthiness focuses mainly on indistinguishable triggers in input space and inseparable backdoor representations in feature space, aiming to circumvent backdoor defenses that examine these respective spaces. However, existing backdoor attacks are typically designed to resist a specific type of backdoor defense without considering the diverse range of defense mechanisms. Based on this observation, we pose a natural question: Are current backdoor attacks truly a real-world threat when facing diverse practical defenses? To answer this question, we examine 12 common backdoor attacks that focus on input-space or feature-space stealthiness and 17 diverse representative defenses. Surprisingly, we reveal a critical blind spot: Backdoor attacks designed to be stealthy in input and feature spaces can be mitigated by examining backdoored models in parameter space. To investigate the underlying causes behind this common vulnerability, we study the characteristics of backdoor attacks in the parameter space. Notably, we find that input- and feature-space attacks introduce prominent backdoor-related neurons in parameter space, which are not thoroughly considered by current backdoor attacks. Taking comprehensive stealthiness into account, we propose a novel supply-chain attack called Grond. Grond limits the parameter changes by a simple yet effective module, Adversarial Backdoor Injection (ABI), which adaptively increases the parameter-space stealthiness during the backdoor injection. Extensive experiments demonstrate that Grond outperforms all 12 backdoor attacks against state-of-the-art (including adaptive) defenses on CIFAR-10, GTSRB, and a subset of ImageNet. In addition, we show that ABI consistently improves the effectiveness of common backdoor attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。