发现视觉编码器中层注意力值是关键漏洞,可生成高效通用对抗扰动。
Exploiting Vision Encoder Vulnerabilities for Universal Adversarial Perturbations on Large Vision-Language Models
- 聚焦视觉编码器中层注意力值,定位高危组件进行攻击
- 单次扰动在多模型上攻击成功率超现有方法,计算开销更低
- 无需文本输入,可跨模型迁移,适合大规模鲁棒性测试
大型视觉语言模型(LVLMs)在多模态任务中表现卓越,但对输入图像中的微小对抗扰动极为敏感。现有攻击多针对视觉编码器的最终输出嵌入,隐含将编码器视为均质攻击面,而对其内部组件的脆弱性系统分析仍不充分。我们证明这种分析至关重要:LVLM视觉编码器的对抗脆弱性呈结构性集中而非均匀分布。基于此,提出视觉编码器脆弱组件定向通用对抗扰动(VEV-UAP),一种任务无关且成本低廉的攻击框架。通过组件与层级的注意力机制分析,我们识别出中层的值组件为关键脆弱点,其显著影响下游语言模型行为。VEV-UAP仅选择性攻击这些组件,生成跨图像共享的单一通用扰动,优化过程不涉及文本输入或语言模型。多个LVLM和任务上的实验表明,VEV-UAP在降低计算开销的同时达到当前最优攻击成功率。此外,一个VEV-UAP可在共享相同视觉编码器的不同语言模型间实现迁移,即使搭配不同语言模型也有效,使其成为可扩展的鲁棒性评估实用框架。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have achieved remarkable performance on multimodal tasks but remain highly vulnerable to small adversarial perturbations in input images. Existing attacks typically target the vision encoder's final output embeddings, implicitly treating the encoder as a uniform attack surface, while a systematic analysis of which internal components are most vulnerable has remained largely unexplored. We show such analysis is essential, as adversarial vulnerability in LVLM vision encoders is structurally concentrated rather than uniformly distributed. Building on this, we propose Vision Encoder Vulnerable-Component-Targeted Universal Adversarial Perturbation (VEV-UAP), a task-agnostic and cost-efficient attack framework. Through a component- and layer-wise analysis of attention mechanisms, we identify the value components in middle layers as critical vulnerabilities that strongly influence downstream language model behavior. VEV-UAP selectively targets these components to generate a single universal perturbation shared across images, without involving textual inputs or the language model during optimization. Experiments across multiple LVLMs and tasks show VEV-UAP achieves state-of-the-art attack success rates with reduced computational overhead. Moreover, a single VEV-UAP transfers across LVLMs sharing the same vision encoder, even when paired with different language models, making it a practical framework for scalable robustness evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。