arXiv:2503.10635cs.CVcs.AI2025-03NeurIPS被引 40

通过局部聚焦扰动,90%以上成功率突破GPT-4等闭源模型防御。

A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1

  • 在关键区域集中扰动,提升语义清晰度以增强攻击效果。
  • 对GPT-4.5/4o/o1等模型攻击成功率超90%,且扰动更小。
  • 方法简单有效,适合研究闭源视觉语言模型安全性的学者。

尽管在开源大视觉语言模型(LVLM)上表现良好,基于迁移的定向攻击常在闭源商业LVLM上失效。分析失败的对抗扰动发现,其通常源自均匀分布且缺乏明确语义细节,导致商业黑盒LVLM要么忽略扰动,要么误读语义,从而攻击失败。为此,我们提出通过在局部区域编码显式语义信息,确保捕捉细粒度特征与跨模型可迁移性,并将修改集中在语义丰富区域而非均匀分布。具体方法为:每轮优化时,随机裁剪图像并控制长宽比与尺度,重缩放后对齐目标图像的嵌入空间。尽管源-目标匹配方法已有先例,我们首次提供严格分析,揭示扰动优化与语义间的紧密联系。实验验证假设:聚焦关键区域的局部聚合扰动具有惊人可迁移性,成功攻击GPT-4.5、GPT-4o、Gemini-2.0-flash、Claude-3.5/3.7-sonnet,甚至推理模型o1、Claude-3.7-thinking与Gemini-2.0-flash-thinking。本方法在GPT-4.5/4o/o1上成功率超90%,显著优于现有最优攻击方法,且$ε_1/ε_2$扰动更小。

原文摘要 · Abstract (English)

Despite promising performance on open-source large vision-language models (LVLMs), transfer-based targeted attacks often fail against closed-source commercial LVLMs. Analyzing failed adversarial perturbations reveals that the learned perturbations typically originate from a uniform distribution and lack clear semantic details, resulting in unintended responses. This critical absence of semantic information leads commercial black-box LVLMs to either ignore the perturbation entirely or misinterpret its embedded semantics, thereby causing the attack to fail. To overcome these issues, we propose to refine semantic clarity by encoding explicit semantic details within local regions, thus ensuring the capture of finer-grained features and inter-model transferability, and by concentrating modifications on semantically rich areas rather than applying them uniformly. To achieve this, we propose a simple yet highly effective baseline: at each optimization step, the adversarial image is cropped randomly by a controlled aspect ratio and scale, resized, and then aligned with the target image in the embedding space. While the naive source-target matching method has been utilized before in the literature, we are the first to provide a tight analysis, which establishes a close connection between perturbation optimization and semantics. Experimental results confirm our hypothesis. Our adversarial examples crafted with local-aggregated perturbations focused on crucial regions exhibit surprisingly good transferability to commercial LVLMs, including GPT-4.5, GPT-4o, Gemini-2.0-flash, Claude-3.5/3.7-sonnet, and even reasoning models like o1, Claude-3.7-thinking and Gemini-2.0-flash-thinking. Our approach achieves success rates exceeding 90% on GPT-4.5, 4o, and o1, significantly outperforming all prior state-of-the-art attack methods with lower $\ell_1/\ell_2$ perturbations.

对抗攻击闭源模型语义聚焦视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。