系统梳理多模态大模型的对抗攻击,揭示其脆弱根源。
Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey
- 按攻击目标分类,统一跨模态攻击框架
- 发现模型脆弱性源于共同的架构与表征缺陷
- 适合安全研究者和模型开发者阅读
多模态大语言模型(MLLMs)融合文本、图像、音频、视频等多种模态信息,实现视觉问答、音频翻译等复杂能力。然而,其更强的表达能力也带来了新的、被放大的对抗攻击风险。本文对MLLMs的对抗威胁进行系统性综述,不仅列举攻击方法,更深入剖析模型易受攻击的根本原因。提出一个基于攻击目标的分类体系,统一不同模态与部署场景下的攻击表面。同时,从漏洞角度分析完整性攻击、安全失效与越狱、指令劫持及训练期投毒等现象,揭示它们共享的架构与表征弱点。该框架为理解MLLMs的对抗行为提供解释基础,并指导构建更鲁棒、更安全的多模态系统。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) integrate information from multiple modalities such as text, images, audio, and video, enabling complex capabilities such as visual question answering and audio translation. While powerful, this increased expressiveness introduces new and amplified vulnerabilities to adversarial manipulation. This survey provides a comprehensive and systematic analysis of adversarial threats to MLLMs, moving beyond enumerating attack techniques to explain the underlying causes of model susceptibility. We introduce a taxonomy that organizes adversarial attacks according to attacker objectives, unifying diverse attack surfaces across modalities and deployment settings. Additionally, we also present a vulnerability-centric analysis that links integrity attacks, safety and jailbreak failures, control and instruction hijacking, and training-time poisoning to shared architectural and representational weaknesses in multimodal systems. Together, this framework provides an explanatory foundation for understanding adversarial behavior in MLLMs and informs the development of more robust and secure multimodal language systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。