通过动态对齐视觉语言连接,提升多模态模型对抗样本的跨模型迁移能力。
Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack
- 在视觉-语言连接处注入动态扰动,增强跨模型泛化性
- 在多个MLLM上实现显著更高的攻击迁移率,包括闭源模型Gemini
- 针对视觉-语言对齐机制设计,适合研究对抗鲁棒性与多模态安全
多模态大语言模型(MLLM)近年来因其图像识别与理解能力受到关注。然而,尽管MLLM易受对抗攻击,其攻击在不同模型间的迁移能力仍有限,尤其在定向攻击场景下。现有方法主要聚焦于视觉模态扰动,难以应对视觉-语言模态对齐的复杂性。本文提出动态视觉-语言对齐(DynVLA)攻击,通过在视觉-语言连接层注入动态扰动,提升对抗样本在不同视觉-语言对齐方式下的通用性。实验结果表明,DynVLA显著提升了对抗样本在多种MLLM(如BLIP2、InstructBLIP、MiniGPT4、LLaVA)及闭源模型Gemini之间的迁移能力。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs), built upon LLMs, have recently gained attention for their capabilities in image recognition and understanding. However, while MLLMs are vulnerable to adversarial attacks, the transferability of these attacks across different models remains limited, especially under targeted attack setting. Existing methods primarily focus on vision-specific perturbations but struggle with the complex nature of vision-language modality alignment. In this work, we introduce the Dynamic Vision-Language Alignment (DynVLA) Attack, a novel approach that injects dynamic perturbations into the vision-language connector to enhance generalization across diverse vision-language alignment of different models. Our experimental results show that DynVLA significantly improves the transferability of adversarial examples across various MLLMs, including BLIP2, InstructBLIP, MiniGPT4, LLaVA, and closed-source models such as Gemini.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。