提出新型动态后门攻击,让视觉语言模型错误定位目标物体。
IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- 用文本引导生成感知隐蔽的动态触发器,精准控制模型输出。
- 在多个数据集和模型上实现超90%攻击成功率,且不降低正常准确率。
- 适合关注多模态模型安全的研究者,揭示视觉定位系统深层漏洞。
近年来,视觉语言模型(VLMs)显著提升了视觉定位任务的性能,该任务基于自然语言查询在图像中定位物体。然而,基于VLM的定位系统安全性尚未得到充分研究。本文揭示了一种新型且现实的漏洞:首个针对VLM视觉定位的多目标后门攻击。不同于依赖静态触发器或固定目标的已有攻击方法,我们提出IAG,通过文本条件化的UNet动态生成输入感知、文本引导的触发器,以特定目标描述为条件执行攻击。该方法将不可察觉的目标语义线索嵌入视觉输入,同时保持良性样本上的正常定位性能。我们进一步设计联合训练目标,平衡语言能力与感知重建,确保触发器的不可察觉性、有效性与隐蔽性。在多个VLM(如LLaVA、InternVL、Ferret)和基准数据集(RefCOCO、RefCOCO+、RefCOCOg、Flickr30k Entities、ShowUI)上的大量实验表明,IAG在几乎所有设置下均达到最优攻击成功率(超过90%),且不影响干净准确率,对现有防御手段具有鲁棒性,并展现出跨数据集与模型的可迁移性。这些发现凸显了具备定位能力的VLM存在重大安全风险,亟需开展可信多模态理解研究。
原文摘要 · Abstract (English)
Recent advances in vision-language models (VLMs) have significantly enhanced the visual grounding task, which involves locating objects in an image based on natural language queries. Despite these advancements, the security of VLM-based grounding systems has not been thoroughly investigated. This paper reveals a novel and realistic vulnerability: the first multi-target backdoor attack on VLM-based visual grounding. Unlike prior attacks that rely on static triggers or fixed targets, we propose IAG, a method that dynamically generates input-aware, text-guided triggers conditioned on any specified target object description to execute the attack. This is achieved through a text-conditioned UNet that embeds imperceptible target semantic cues into visual inputs while preserving normal grounding performance on benign samples. We further develop a joint training objective that balances language capability with perceptual reconstruction to ensure imperceptibility, effectiveness, and stealth. Extensive experiments on multiple VLMs (e.g., LLaVA, InternVL, Ferret) and benchmarks (RefCOCO, RefCOCO+, RefCOCOg, Flickr30k Entities, and ShowUI) demonstrate that IAG achieves the best ASRs compared with other baselines on almost all settings without compromising clean accuracy, maintaining robustness against existing defenses, and exhibiting transferability across datasets and models. These findings underscore critical security risks in grounding-capable VLMs and highlight the need for further research on trustworthy multimodal understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。