提出新型视觉变压器后门攻击,任意位置触发仍高效隐蔽。
PASTA: A Patch-Agnostic Twofold-Stealthy Backdoor Attack on Vision Transformers

- 通过多位置触发插入增强跨块注意力辐射效应
- 平均攻击成功率99.13%,视觉与注意力隐蔽性提升144倍以上
- 适合研究模型安全与防御的学者参考
视觉变换器(ViTs)在视觉任务中表现卓越,但易受后门攻击。现有基于块的攻击通常假设推理时触发器位于固定位置以最大化注意力,却忽略了ViT的自注意力机制对长距离块依赖的捕捉能力。本文观察到,当触发器激活邻近块时可实现高攻击效果,这一现象称为触发辐射效应(TRE)。进一步发现,训练时跨块插入触发器可协同增强TRE。此前针对ViT的攻击常牺牲视觉和注意力隐蔽性,易被检测。基于此,我们提出PASTA,一种在像素与注意力双域均隐蔽的两阶段后门攻击。该方法支持在任意块位置触发,引入多位置触发策略增强TRE。然而,在保持隐蔽性的前提下维持强TRE极具挑战,因隐蔽约束会削弱TRE。为此,我们构建双层优化问题,提出自适应后门学习框架,使模型与触发器迭代适应,避免局部最优。大量实验表明,PASTA在四个数据集上平均攻击成功率达99.13%,同时显著提升视觉(144.43倍)与注意力(18.68倍)隐蔽性,并在面对先进ViT防御时具备2.79倍更强鲁棒性,优于基于CNN与ViT的基线方法。
原文摘要 · Abstract (English)
Vision Transformers (ViTs) have achieved remarkable success across vision tasks, yet recent studies show they remain vulnerable to backdoor attacks. Existing patch-wise attacks typically assume a single fixed trigger location during inference to maximize trigger attention. However, they overlook the self-attention mechanism in ViTs, which captures long-range dependencies across patches. In this work, we observe that a patch-wise trigger can achieve high attack effectiveness when activating backdoors across neighboring patches, a phenomenon we term the Trigger Radiating Effect (TRE). We further find that inter-patch trigger insertion during training can synergistically enhance TRE compared to single-patch insertion. Prior ViT-specific attacks that maximize trigger attention often sacrifice visual and attention stealthiness, making them detectable. Based on these insights, we propose PASTA, a twofold stealthy patch-wise backdoor attack in both pixel and attention domains. PASTA enables backdoor activation when the trigger is placed at arbitrary patches during inference. To achieve this, we introduce a multi-location trigger insertion strategy to enhance TRE. However, preserving stealthiness while maintaining strong TRE is challenging, as TRE is weakened under stealthy constraints. We therefore formulate a bi-level optimization problem and propose an adaptive backdoor learning framework, where the model and trigger iteratively adapt to each other to avoid local optima. Extensive experiments show that PASTA achieves 99.13% attack success rate across arbitrary patches on average, while significantly improving visual and attention stealthiness (144.43x and 18.68x) and robustness (2.79x) against state-of-the-art ViT defenses across four datasets, outperforming CNN- and ViT-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。