arXiv:2506.07214cs.CVcs.CR2025-06被引 15

通过语义错配隐蔽植入后门,让视觉语言模型在特定条件下被操控。

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation

  • 利用图像与文本语义错配作为隐性触发器,实现数据投毒攻击。
  • 在4个主流VLM上平均攻击成功率超98%,且跨数据集泛化能力强。
  • 攻击隐蔽性强,现有防御策略均无效,凸显语义漏洞风险。

视觉语言模型(VLMs)虽表现优异,却易受后门攻击影响,攻击者可通过隐藏触发器操纵模型输出。现有攻击多依赖单模态触发器,未充分探索VLM的跨模态融合特性。本文提出新攻击面:利用跨模态语义不一致作为隐性触发器,设计BadSem(基于语义操控的后门攻击),通过训练时故意错位图像-文本对注入隐蔽后门。构建专用数据集SIMBad,用于颜色与物体属性的语义操控。在四个主流VLM上的实验表明,BadSem平均攻击成功率超过98%,对分布外数据泛化良好,并能跨模态迁移。注意力可视化分析显示,模型在语义错配时聚焦敏感区域,而正常输入下行为不变。尝试两种防御策略——系统提示和监督微调,均未能有效缓解该攻击。研究揭示了VLM语义层面的严重安全隐患,亟需重视其安全部署。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have shown remarkable performance, but are also vulnerable to backdoor attacks whereby the adversary can manipulate the model's outputs through hidden triggers. Prior attacks primarily rely on single-modality triggers, leaving the crucial cross-modal fusion nature of VLMs largely unexplored. Unlike prior work, we identify a novel attack surface that leverages cross-modal semantic mismatches as implicit triggers. Based on this insight, we propose BadSem (Backdoor Attack with Semantic Manipulation), a data poisoning attack that injects stealthy backdoors by deliberately misaligning image-text pairs during training. To perform the attack, we construct SIMBad, a dataset tailored for semantic manipulation involving color and object attributes. Extensive experiments across four widely used VLMs show that BadSem achieves over 98% average ASR, generalizes well to out-of-distribution datasets, and can transfer across poisoning modalities. Our detailed analysis using attention visualization shows that backdoored models focus on semantically sensitive regions under mismatched conditions while maintaining normal behavior on clean inputs. To mitigate the attack, we try two defense strategies based on system prompt and supervised fine-tuning but find that both of them fail to mitigate the semantic backdoor. Our findings highlight the urgent need to address semantic vulnerabilities in VLMs for their safer deployment.

后门攻击视觉语言模型语义安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。