arXiv:2503.13962cs.CV2025-03综述被引 18

系统梳理多模态大模型的对抗脆弱性,揭示跨模态攻击风险。

Survey of Adversarial Robustness in Multimodal Large Language Models

  • 按模态分类梳理对抗攻击类型与机制
  • 总结主流评测数据集与评估指标体系
  • 适合关注AI安全与多模态模型可信性的研究者

多模态大语言模型(MLLMs)在融合文本、图像、视频、音频和语音等多种模态理解方面表现出色。然而,其在实际应用中的部署引发了对抗脆弱性的重大担忧,可能影响模型的安全性与可靠性。与单模态模型不同,MLLMs因模态间的相互依赖,易受模态特定威胁及跨模态对抗干扰。本文系统回顾了MLLMs的对抗鲁棒性,涵盖不同模态。首先介绍MLLMs基础并构建针对性的对抗攻击分类体系;随后综述关键数据集与评估指标;进一步深入分析各模态下的攻击方法;最后指出核心挑战并提出未来研究方向。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated exceptional performance in artificial intelligence by facilitating integrated understanding across diverse modalities, including text, images, video, audio, and speech. However, their deployment in real-world applications raises significant concerns about adversarial vulnerabilities that could compromise their safety and reliability. Unlike unimodal models, MLLMs face unique challenges due to the interdependencies among modalities, making them susceptible to modality-specific threats and cross-modal adversarial manipulations. This paper reviews the adversarial robustness of MLLMs, covering different modalities. We begin with an overview of MLLMs and a taxonomy of adversarial attacks tailored to each modality. Next, we review key datasets and evaluation metrics used to assess the robustness of MLLMs. After that, we provide an in-depth review of attacks targeting MLLMs across different modalities. Our survey also identifies critical challenges and suggests promising future research directions.

多模态对抗鲁棒性AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。