arXiv:2604.04630cs.CV2026-04被引 1

用涂鸦和跨语言文本构造隐蔽后门,攻击自动驾驶视觉语言模型

Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers

  • 采用涂鸦图案和跨语言文本作为自然触发器,隐蔽性强
  • 仅10%污染率即达90%攻击成功率,误报率为0
  • 攻击不降低正常性能,甚至提升指标,难被传统方法发现

视觉语言模型(VLM)正快速融入自动驾驶等安全关键系统,成为潜在后门攻击的靶点。现有攻击多依赖单模态、显式且易检测的触发器,在自动驾驶场景中难以构建隐蔽稳定的攻击通道。本文提出GLA,引入两种自然触发器:通过稳定扩散修复生成的涂鸦视觉模式,可无缝融入城市环境;以及跨语言文本触发器,在保持语义一致的同时引入分布偏移,增强语言侧触发信号的鲁棒性。在DriveVLM上的实验表明,仅需10%的污染比例即可实现90%的攻击成功率(ASR)和0%的误报率(FPR)。更隐蔽的是,该后门不会削弱模型在干净样本上的性能,反而提升如BLEU-1等指标,使基于性能下降的传统检测方法失效。本研究揭示了自动驾驶VLM中被低估的安全风险,并为安全关键多模态系统中的后门评估提供了新范式。

原文摘要 · Abstract (English)

Visual language model (VLM) is rapidly being integrated into safety-critical systems such as autonomous driving, making it an important attack surface for potential backdoor attacks. Existing backdoor attacks mainly rely on unimodal, explicit, and easily detectable triggers, making it difficult to construct both covert and stable attack channels in autonomous driving scenarios. GLA introduces two naturalistic triggers: graffiti-based visual patterns generated via stable diffusion inpainting, which seamlessly blend into urban scenes, and cross-language text triggers, which introduce distributional shifts while maintaining semantic consistency to build robust language-side trigger signals. Experiments on DriveVLM show that GLA requires only a 10\% poisoning ratio to achieve a 90\% Attack Success Rate (ASR) and a 0\% False Positive Rate (FPR). More insidiously, the backdoor does not weaken the model on clean tasks, but instead improves metrics such as BLEU-1, making it difficult for traditional performance-degradation-based detection methods to identify the attack. This study reveals underestimated security threats in self-driving VLMs and provides a new attack paradigm for backdoor evaluation in safety-critical multimodal systems.

后门攻击自动驾驶多模态VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。