arXiv:2608.07535cs.LGcs.AI2026-08中稿 · ICLR综述

多模态大模型安全风险升级,本文系统梳理新威胁与防护策略。

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

论文配图:Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
图 1 · 摘自论文原文
  • 构建多模态安全威胁分类体系,揭示跨模态风险机制。
  • 发现现有单模态安全框架无法覆盖多模态融合带来的新风险。
  • 适合关注AI安全、多模态系统设计的研究者与工程师阅读。

多模态大语言模型(MLLMs)通过模态对齐与融合整合异构信息,提升理解与推理能力。然而,这种架构变革重塑了机器学习的安全格局。模型复杂度增加及跨模态交互催生新型威胁,包括模态融合受损、模态错位以及融合后的安全风险,反映出超越单模态假设的威胁模型演变。这些变化对安全解决方案提出新约束,现有基于单模态学习的框架难以涵盖。为此,本文系统分析MLLMs的安全演进态势:首先提出多模态基础的安全威胁分类体系,分析威胁模型的转变,涵盖对抗攻击、数据投毒、越狱和幻觉;其次总结更新的安全假设,并归类近期的MLLM安全策略进展;最后探讨开放挑战与未来方向,以推动更严谨、可扩展的多模态系统安全机制发展。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of machine learning. Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fused safety risks, reflecting shifts in threat modeling beyond uni-modal assumptions. These shifts, in turn, impose new constraints on safety solutions not captured by existing frameworks rooted in uni-modal learning. Motivated by these challenges, this survey provides a systematic analysis of the evolving safety landscape of MLLMs. We first propose a multimodal grounded taxonomy of safety threats and analyze shifts in threat models, covering adversarial attacks, data poisoning, jailbreaks, and hallucinations. We then summarize updated safety assumptions and organize recent advances in MLLM safety strategies accordingly. Finally, we discuss open challenges and future directions to inform the development of more principled and scalable safety mechanisms for multimodal systems.

多模态安全大模型威胁建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。