arXiv:2506.05429cs.CVcs.AI2025-06CVPR

提出协同扰动框架,评估视觉语言模型在双模态攻击下的鲁棒性。

Coordinated Robustness Evaluation Framework for Vision-Language Models

  • 训练通用代理模型生成图像与文本的联合表示,实现双模态协同攻击。
  • 在VQA和视觉推理数据集上,攻击成功率显著高于单模态与现有多模态方法。
  • 适用于评估Instruct-BLIP、ViLT等主流视觉语言模型的脆弱性。

视觉语言模型融合了计算机视觉与自然语言处理能力,在图像描述和视觉问答等任务中表现突出。然而,与传统模型类似,它们对微小扰动敏感,部署场景下鲁棒性面临挑战。评估其鲁棒性需同时在视觉与语言模态施加扰动,以捕捉跨模态依赖关系。本文训练一个通用代理模型,可同时接收图像与文本输入,生成联合表示,并据此生成针对图像与文本的对抗扰动。该协同攻击策略在视觉问答与视觉推理数据集上,针对多种前沿视觉语言模型进行评估。结果表明,该方法优于现有单模态及多模态攻击策略,有效暴露了Instruct-BLIP、ViLT等先进预训练多模态模型的脆弱性。

原文摘要 · Abstract (English)

Vision-language models, which integrate computer vision and natural language processing capabilities, have demonstrated significant advancements in tasks such as image captioning and visual question and answering. However, similar to traditional models, they are susceptible to small perturbations, posing a challenge to their robustness, particularly in deployment scenarios. Evaluating the robustness of these models requires perturbations in both the vision and language modalities to learn their inter-modal dependencies. In this work, we train a generic surrogate model that can take both image and text as input and generate joint representation which is further used to generate adversarial perturbations for both the text and image modalities. This coordinated attack strategy is evaluated on the visual question and answering and visual reasoning datasets using various state-of-the-art vision-language models. Our results indicate that the proposed strategy outperforms other multi-modal attacks and single-modality attacks from the recent literature. Our results demonstrate their effectiveness in compromising the robustness of several state-of-the-art pre-trained multi-modal models such as instruct-BLIP, ViLT and others.

多模态鲁棒性对抗攻击视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。