arXiv:2605.26501cs.CVcs.AI2026-05AAAI被引 23

提出可同时扰动图像和文本的通用黑盒攻击,揭露视觉语言模型脆弱性。

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

论文配图:Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization
图 1 · 摘自论文原文
  • 联合优化图像纹理扰动与文本提示扰动,实现跨模态协同攻击。
  • 仅用模型查询即可生成强迁移性攻击,在主流视觉语言模型上成功率超85%。
  • 适合安全研究者评估模型鲁棒性,尤其关注多模态系统风险的团队。

大规模视觉语言模型(LVLMs)在图像描述、视觉问答等任务中表现优异,但其对多模态对抗攻击的鲁棒性尚未充分探索,可能危及自动驾驶、内容审核等关键应用。现有攻击多局限于单模态或需白盒访问,实用性受限。本文提出多模态对抗协同框架(MMAS),可生成无需白盒信息的通用黑盒攻击。该方法同步生成图像上的纹理约束型通用对抗扰动(基于小波变换确保不可察觉性)和文本上的嵌入空间L-norm约束提示扰动(保持语义连贯性),并通过新提出的跨模态正则项对齐扰动梯度方向,增强协同效应与迁移能力。实验表明,该攻击在多个主流LVLM上具有强通用性,成功率超过85%。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial attacks, particularly those exploiting both modalities, remains underexplored, posing risks to critical applications like autonomous driving and content moderation. Existing attacks focus on single modalities or require impractical white-box access, limiting their real-world relevance. In this paper, we introduce Multi-Modal Adversarial Synergy, a groundbreaking framework that crafts universal, black-box multi-modal attacks against LVLMs. MMAS simultaneously generates a texture scale-constrained universal adversarial perturbation for images and a learnable prompt perturbation for text, optimized jointly using only model queries. The image perturbation leverages wavelet-based texture constraints to ensure imperceptibility and robustness across diverse visual inputs. The text perturbation, constrained by an L-norm in the embedding space, maintains semantic coherence while steering outputs toward a target. A novel cross-modal regularization term aligns the perturbations' gradient directions, enhancing their synergistic impact and transferability across tasks and models. Extensive experiments show the strong universal adversarial capabilities of our proposed attack with prevalent LVLMs.

多模态攻击对抗样本视觉语言模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。