arXiv:2603.16960cs.CRcs.AI2026-03被引 1

测试开源视觉语言模型在电商环境下的抗攻击能力,发现攻击成功率超66%。

Adversarial attacks against Modern Vision-Language Models

  • 用三种梯度攻击法测试模型鲁棒性
  • LLaVA攻击成功率达66.9%,Qwen更抗攻击
  • 结果对模型上线前安全评估有直接指导意义

我们研究了在自包含电商环境中部署的开源视觉语言模型(VLM)代理的对抗鲁棒性,该环境模拟真实预部署条件。评估了两个代理:LLaVA-v1.5-7B 和 Qwen2.5-VL-7B,采用三种基于梯度的攻击方法:基本迭代法(BIM)、投影梯度下降(PGD)和基于CLIP的频谱攻击。针对LLaVA,三种攻击的成功率分别为52.6%、53.8%和66.9%,表明简单梯度方法对开源VLM代理构成实际威胁。Qwen2.5-VL在所有攻击下表现显著更鲁棒(成功率分别为6.5%、7.7%和15.5%),说明不同开源VLM系列在对抗韧性上存在显著架构差异。这些发现对VLM代理商业部署前的安全评估具有直接意义。

原文摘要 · Abstract (English)

We study adversarial robustness of open-source vision-language model (VLM) agents deployed in a self-contained e-commerce environment built to simulate realistic pre-deployment conditions. We evaluate two agents, LLaVA-v1.5-7B and Qwen2.5-VL-7B, under three gradient-based attacks: the Basic Iterative Method (BIM), Projected Gradient Descent (PGD), and a CLIP-based spectral attack. Against LLaVA, all three attacks achieve substantial attack success rates (52.6%, 53.8%, and 66.9% respectively), demonstrating that simple gradient-based methods pose a practical threat to open-source VLM agents. Qwen2.5-VL proves significantly more robust across all attacks (6.5%, 7.7%, and 15.5%), suggesting meaningful architectural differences in adversarial resilience between open-source VLM families. These findings have direct implications for the security evaluation of VLM agents prior to commercial deployment.

视觉语言模型对抗攻击安全评估LLaVA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。