arXiv:2411.15720cs.CV2024-11CVPR被引 40

提出链式攻击方法,提升视觉语言模型的对抗攻击效果。

Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks

  • 通过多模态语义迭代更新生成更强对抗样本
  • 黑盒攻击下实现高成功率且无需目标模型信息
  • 为评估视觉语言模型安全提供新方法,适合安全研究者

预训练视觉语言模型(VLMs)在图像理解与自然语言生成任务中表现卓越,如图像描述和响应生成。随着其应用日益广泛,安全与鲁棒性问题引发关注,攻击者可能通过恶意攻击诱导模型生成有害内容。现有基于迁移的黑盒攻击常忽略视觉与文本模态间的语义关联,导致对抗样本生成效果不佳。为此,本文提出链式攻击(Chain of Attack, CoA),通过一系列中间攻击步骤,迭代优化多模态语义对齐,显著提升对抗样本的迁移性和攻击效率。同时,提出统一的攻击成功率计算方法,支持自动逃逸评估。在最真实、高风险场景下的大量实验表明,CoA仅需黑盒攻击即可有效诱导模型生成目标响应,且无需了解目标模型结构。本研究全面揭示了VLMs的脆弱性,为未来模型安全性设计提供重要参考。

原文摘要 · Abstract (English)

Pre-trained vision-language models (VLMs) have showcased remarkable performance in image and natural language understanding, such as image captioning and response generation. As the practical applications of vision-language models become increasingly widespread, their potential safety and robustness issues raise concerns that adversaries may evade the system and cause these models to generate toxic content through malicious attacks. Therefore, evaluating the robustness of open-source VLMs against adversarial attacks has garnered growing attention, with transfer-based attacks as a representative black-box attacking strategy. However, most existing transfer-based attacks neglect the importance of the semantic correlations between vision and text modalities, leading to sub-optimal adversarial example generation and attack performance. To address this issue, we present Chain of Attack (CoA), which iteratively enhances the generation of adversarial examples based on the multi-modal semantic update using a series of intermediate attacking steps, achieving superior adversarial transferability and efficiency. A unified attack success rate computing method is further proposed for automatic evasion evaluation. Extensive experiments conducted under the most realistic and high-stakes scenario, demonstrate that our attacking strategy can effectively mislead models to generate targeted responses using only black-box attacks without any knowledge of the victim models. The comprehensive robustness evaluation in our paper provides insight into the vulnerabilities of VLMs and offers a reference for the safety considerations of future model developments.

对抗攻击视觉语言模型黑盒攻击安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。