arXiv:2506.22982cs.CV2025-06中稿 · MLRC 2025

提升视觉语言模型的跨提示对抗攻击效果,增强安全性研究。

Revisiting CroPA: A Reproducibility Study and Enhancements for Cross-Prompt Adversarial Transferability in Vision-Language Models

  • 提出新初始化策略,显著提高攻击成功率。
  • 通过通用扰动实现跨图像攻击,扩展攻击范围。
  • 设计关注视觉编码器注意力的新损失函数,提升泛化能力。

大型视觉语言模型(VLMs)已革新计算机视觉,支持图像分类、图文生成和视觉问答等任务。然而,它们在视觉与文本模态均可被操纵的场景下仍易受对抗攻击影响。本文对《一张图胜过千言:视觉语言模型中跨提示的对抗可迁移性》进行复现研究,验证了跨提示攻击(CroPA)的优越性,并在此基础上提出多项改进:(1) 一种新初始化策略,显著提升攻击成功率(ASR);(2) 通过学习通用扰动,探索跨图像可迁移性;(3) 设计针对视觉编码器注意力机制的新损失函数,提升泛化能力。在Flamingo、BLIP-2、InstructBLIP及扩展实验中的LLaVA上评估表明,原结果可复现,且改进方法持续增强对抗有效性。本工作强化了对VLMs对抗脆弱性的研究,为生成更具可迁移性的对抗样本提供了更稳健框架,对理解VLM在真实应用中的安全性具有重要意义。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) have revolutionized computer vision, enabling tasks such as image classification, captioning, and visual question answering. However, they remain highly vulnerable to adversarial attacks, particularly in scenarios where both visual and textual modalities can be manipulated. In this study, we conduct a comprehensive reproducibility study of "An Image is Worth 1000 Lies: Adversarial Transferability Across Prompts on Vision-Language Models" validating the Cross-Prompt Attack (CroPA) and confirming its superior cross-prompt transferability compared to existing baselines. Beyond replication we propose several key improvements: (1) A novel initialization strategy that significantly improves Attack Success Rate (ASR). (2) Investigate cross-image transferability by learning universal perturbations. (3) A novel loss function targeting vision encoder attention mechanisms to improve generalization. Our evaluation across prominent VLMs -- including Flamingo, BLIP-2, and InstructBLIP as well as extended experiments on LLaVA validates the original results and demonstrates that our improvements consistently boost adversarial effectiveness. Our work reinforces the importance of studying adversarial vulnerabilities in VLMs and provides a more robust framework for generating transferable adversarial examples, with significant implications for understanding the security of VLMs in real-world applications.

对抗攻击视觉语言模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。