arXiv:2511.20494cs.CL2025-11被引 2

用对抗扰动让多模态大模型输出混乱,破坏AI代理运行

Adversarial Confusion Attack: Disrupting Multimodal Large Language Models

  • 用多个开源模型联合最大化下一个词的熵来生成干扰图像
  • 一张对抗图像可使同组模型在全图和验证码场景下失效
  • 扰动能跨模型迁移,影响开源与闭源多模态模型

我们提出一种新型威胁——对抗混淆攻击,针对多模态大语言模型(MLLMs)。与越狱或定向误分类不同,其目标是引发系统性干扰,使模型生成不连贯或自信错误的输出。实际应用场景包括将对抗图像嵌入网页,以阻止基于MLLM的AI智能体可靠运行。该攻击通过小规模开源MLLM集成,利用基本的PGD对抗技术最大化下一词熵。在白盒设置下,单张对抗图像即可破坏集成中所有模型,无论是在全图输入还是对抗验证码场景。尽管使用基础方法,生成的扰动仍可有效迁移到未见过的开源模型(如Qwen3-VL)及闭源模型(如GPT-5.1)。

原文摘要 · Abstract (English)

We introduce the Adversarial Confusion Attack, a new class of threats against multimodal large language models (MLLMs). Unlike jailbreaks or targeted misclassification, the goal is to induce systematic disruption that makes the model generate incoherent or confidently incorrect outputs. Practical applications include embedding such adversarial images into websites to prevent MLLM-powered AI Agents from operating reliably. The proposed attack maximizes next-token entropy using a small ensemble of open-source MLLMs. In the white-box setting, we show that a single adversarial image can disrupt all models in the ensemble, both in the full-image and Adversarial CAPTCHA settings. Despite relying on a basic adversarial technique (PGD), the attack generates perturbations that transfer to both unseen open-source (e.g., Qwen3-VL) and proprietary (e.g., GPT-5.1) models.

对抗攻击多模态大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。