arXiv:2602.09431cs.CRcs.CV2026-02

针对视觉语言模型的对抗攻击,提出基于文本定位证据的新方法。

Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models

  • 通过文本引导的视觉区域分配与干扰,优化对抗扰动
  • 在多种模型上实现更高黑盒迁移成功率,超越现有方法
  • 适合研究模型鲁棒性与防御机制的学者参考

大型视觉语言模型(LVLM)在多模态任务中表现优异,但其对视觉输入的依赖使其易受对抗攻击。基于编码器的攻击通过仅优化视觉编码器来生成扰动,效率较高。然而,现有方法通常假设代理编码器与目标模型的视觉编码器相同或相似。本文系统研究了在异构架构下的黑盒部署中攻击的迁移能力,发现不同模型间特定视觉证据不一致,而文本条件下的定位区域更稳定且与图像描述相关。现有攻击对这些区域对齐不足、干扰不充分。为此,提出接地驱动攻击(GDA),将扰动优化聚焦于文本引导的证据区域:通过感知接地的扰动分配集中预算,通过以接地为中心的证据破坏增强全局与局部干扰。在多个目标模型和任务上的实验表明,GDA在黑盒迁移中持续优于现有编码器级攻击。结果强调了文本引导证据在对抗迁移中的核心作用,推动了面向接地的鲁棒性评估与防御设计。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have achieved impressive performance across multimodal tasks, but their reliance on visual inputs exposes them to adversarial threats. Encoder-based attacks provide an efficient alternative to end-to-end optimization by crafting perturbations through the vision encoder alone. However, existing encoder-based attacks often assume that the surrogate encoder is identical or similar to the victim LVLM's vision encoder. In this work, we present a systematic study of their transferability in more realistic black-box deployments with heterogeneous LVLM architectures. We find that model-specific visual evidence is inconsistent across models, whereas text-conditioned grounding regions are more closely tied to caption-relevant evidence and provide a more stable transfer target. However, existing attacks remain weakly aligned with and insufficiently disrupt these regions. Motivated by these findings, we propose Grounding-Driven Attack (GDA), which aligns perturbation optimization with text-grounded evidence. GDA combines Grounding-Aware Perturbation Allocation to concentrate perturbation budget on grounded evidence regions with Grounding-Centric Evidence Disruption to intensify their global and local disruption. Experiments across diverse victim models and tasks show that GDA consistently outperforms existing encoder-based attacks in black-box transfer. These results highlight the central role of text-grounded evidence in adversarial transferability and motivate grounding-aware robustness evaluation and defense design.

对抗攻击视觉语言模型文本引导鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。