提出通用对抗攻击方法,让单一扰动可骗过多个闭源多模态模型。
Universal Adversarial Attacks against Closed-Source MLLMs via Target-View Routed Meta Optimization
- 通过多裁剪聚合与注意力引导,稳定攻击目标监督信号。
- 在未知图像上对GPT-4o和Gemini-2.0的攻击成功率分别提升23.7%和19.9%。
- 适合研究模型安全、对抗攻击的学者与工程师使用。
针对闭源多模态大语言模型(MLLMs)的目标性对抗攻击在黑盒迁移场景中日益受到关注,但现有方法多为样本特异性,跨输入复用性差。本文研究更严格的设定——通用目标可迁移对抗攻击(UTTAA),即单个扰动需持续将任意输入引导至指定目标,适用于未知的商用MLLMs。直接将现有样本级攻击推广至通用设置面临三大挑战:(i) 目标裁剪随机性导致目标监督方差高;(ii) 通用性抑制图像特有线索,使词元对齐不可靠;(iii) 每目标少样本适应对初始化高度敏感,影响性能。为此,我们提出MCRMO-Attack,通过多裁剪聚合与注意力引导的裁剪选择稳定监督信号,利用可对齐性门控的词元路由提升词元级可靠性,并元学习跨目标扰动先验以生成更强的每目标解。在多个商用MLLM上,相较最强通用基线,对未见图像的攻击成功率在GPT-4o上提升23.7%,在Gemini-2.0上提升19.9%。
原文摘要 · Abstract (English)
Targeted adversarial attacks on closed-source multimodal large language models (MLLMs) have been increasingly explored under black-box transfer, yet prior methods are predominantly sample-specific and offer limited reusability across inputs. We instead study a more stringent setting, Universal Targeted Transferable Adversarial Attacks (UTTAA), where a single perturbation must consistently steer arbitrary inputs toward a specified target across unknown commercial MLLMs. Naively adapting existing sample-wise attacks to this universal setting faces three core difficulties: (i) target supervision becomes high-variance due to target-crop randomness, (ii) token-wise matching is unreliable because universality suppresses image-specific cues that would otherwise anchor alignment, and (iii) few-source per-target adaptation is highly initialization-sensitive, which can degrade the attainable performance. In this work, we propose MCRMO-Attack, which stabilizes supervision via Multi-Crop Aggregation with an Attention-Guided Crop, improves token-level reliability through alignability-gated Token Routing, and meta-learns a cross-target perturbation prior that yields stronger per-target solutions. Across commercial MLLMs, we boost unseen-image attack success rate by +23.7\% on GPT-4o and +19.9\% on Gemini-2.0 over the strongest universal baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。