针对视觉状态空间模型设计了隐蔽性强的后门攻击方法。
BadViM: Backdoor Attack against Vision Mamba
- 利用频率敏感性设计分布式触发器,实现隐蔽攻击。
- 通过隐藏状态对齐损失提升攻击成功率,达98.7%以上。
- 可抵抗多种防御手段,适合研究模型安全的人员。
视觉状态空间模型(SSMs),尤其是视觉马尔可夫模型(ViM),正成为视觉变压器(ViTs)的有力替代方案。然而,该新型架构的安全性,特别是其对后门攻击的脆弱性,仍严重缺乏研究。后门攻击旨在向目标模型植入隐藏触发器,使模型在遇到含触发器输入时误分类,同时保持对干净输入的正常表现。本文提出针对视觉马尔可夫模型的新型后门攻击框架BadViM,其核心为共振频率触发器(RFT),利用受害模型的频率敏感特性生成隐蔽、分布式的触发器。为最大化攻击效果,引入隐藏状态对齐损失,通过调整后门图像与目标类别隐藏状态的一致性来操控内部表示。大量实验表明,BadViM在保持高清洁数据准确率的同时,实现了超过98.7%的攻击成功率。此外,该攻击对常见防御措施如PatchDrop、PatchShuffle和JPEG压缩表现出显著鲁棒性,通常能中和普通后门攻击。
原文摘要 · Abstract (English)
Vision State Space Models (SSMs), particularly architectures like Vision Mamba (ViM), have emerged as promising alternatives to Vision Transformers (ViTs). However, the security implications of this novel architecture, especially their vulnerability to backdoor attacks, remain critically underexplored. Backdoor attacks aim to embed hidden triggers into victim models, causing the model to misclassify inputs containing these triggers while maintaining normal behavior on clean inputs. This paper investigates the susceptibility of ViM to backdoor attacks by introducing BadViM, a novel backdoor attack framework specifically designed for Vision Mamba. The proposed BadViM leverages a Resonant Frequency Trigger (RFT) that exploits the frequency sensitivity patterns of the victim model to create stealthy, distributed triggers. To maximize attack efficacy, we propose a Hidden State Alignment loss that strategically manipulates the internal representations of model by aligning the hidden states of backdoor images with those of target classes. Extensive experimental results demonstrate that BadViM achieves superior attack success rates while maintaining clean data accuracy. Meanwhile, BadViM exhibits remarkable resilience against common defensive measures, including PatchDrop, PatchShuffle and JPEG compression, which typically neutralize normal backdoor attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。