仅扰动视觉编码器就能骗过视觉语言模型,揭示其安全短板
Breaking the weakest link to evade vision language models

- 只优化视觉编码器生成对抗样本,降低计算开销
- 微小不可察觉的图像扰动能大幅改变模型输出文本
- 适合关注多模态模型安全性的研究者与开发者
视觉语言模型(VLMs)已成为多模态AI系统的关键组件,支持在现实和安全敏感场景中对视觉与文本输入进行联合推理。尽管部署日益广泛,其对对抗性威胁的鲁棒性仍缺乏充分研究,尤其在针对多模态对齐的逃避攻击方面。本文研究了对视觉输入施加对抗扰动时VLMs的脆弱性,考察两类攻击:非目标攻击(破坏原图语义理解)与目标攻击(迫使模型生成无关的特定描述)。我们提出一种基于梯度的攻击方法,仅在VLM的视觉编码器上进行优化,而非整个多模态架构,显著降低计算成本与资源需求,同时保持强攻击效果。我们在Qwen2.5-VL、Granite-Vision、FastVLM和Phi-3.5-Vision等开源VLM上验证该方法,结果表明,微小且人眼难以察觉的扰动即可显著改变模型输出的文本解释。研究揭示现代VLMs在对抗操纵下的脆弱性,强调需加强多模态AI系统的鲁棒性与安全机制。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) have recently emerged as a critical component of multimodal AI systems, enabling joint reasoning over visual and textual inputs in real-world and safety-critical applications. Despite their growing deployment, the robustness of VLMs against adversarial threats remains insufficiently explored, particularly in the context of evasion attacks targeting multimodal alignment. In this work, we investigate the vulnerability of VLMs to adversarial perturbations applied to visual inputs and study two attack settings: untargeted attacks, where the goal is to disrupt the model's interpretation of the original image, and targeted attacks, where the adversary aims to force the model to generate a specific semantic description unrelated to the original image. To efficiently generate adversarial examples, we propose a gradient-based attack method that performs optimization exclusively on the vision encoder of the VLM rather than on the entire multimodal architecture. This design significantly reduces the computational cost and resource requirements of the attack while maintaining strong effectiveness. We evaluate our approach on several open-source VLMs, including Qwen2.5-VL, Granite-Vision, FastVLM, and Phi-3.5-Vision, and show that small, human-imperceptible perturbations can substantially alter the textual interpretation produced by the models. Our findings highlight the vulnerability of modern VLMs to adversarial manipulation and emphasize the need for improved robustness and security mechanisms in multimodal AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。