提出黑盒通用扰动攻击多图像大模型,提升攻击成功率。
LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models
- 基于注意力约束,让模型无法有效融合多图信息。
- 跨图传染机制使扰动扩散,无需修改全部输入。
- 位置无关攻击策略增强鲁棒性,适合真实场景。
多模态大语言模型(MLLMs)在视觉-语言任务中表现优异,可处理多图像输入。然而,多图像MLLM的脆弱性尚未被充分研究。现有对抗攻击主要针对单图像场景,且常假设白盒威胁模型,在现实场景中不适用。本文提出LAMP,一种针对多图像MLLM的黑盒通用对抗扰动方法。LAMP引入基于注意力的约束,阻止模型有效聚合多图信息;设计新颖的跨图传染约束,使扰动能影响干净特征,实现无须全量修改的传播式攻击;并采用索引-注意力抑制损失,实现位置无关的鲁棒攻击。实验表明,LAMP超越当前最优基线,在多个视觉-语言任务与模型上均取得最高攻击成功率。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable performance across vision-language tasks. Recent advancements allow these models to process multiple images as inputs. However, the vulnerabilities of multi-image MLLMs remain unexplored. Existing adversarial attacks focus on single-image settings and often assume a white-box threat model, which is impractical in many real-world scenarios. This paper introduces LAMP, a black-box method for learning Universal Adversarial Perturbations (UAPs) targeting multi-image MLLMs. LAMP applies an attention-based constraint that prevents the model from effectively aggregating information across images. LAMP also introduces a novel cross-image contagious constraint that forces perturbed tokens to influence clean tokens, spreading adversarial effects without requiring all inputs to be modified. Additionally, an index-attention suppression loss enables a robust position-invariant attack. Experimental results show that LAMP outperforms SOTA baselines and achieves the highest attack success rates across multiple vision-language tasks and models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。