arXiv:2411.09259cs.CVcs.CL2024-11综述被引 39

系统梳理多模态生成模型的越狱攻击与防御方法,助你安全使用AI。

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey

  • 从输入到输出四层分析越狱攻击路径,覆盖文本、图像等多模态场景。
  • 提出攻击与防御的详细分类体系,涵盖任意模态间生成任务。
  • 适合关注AI安全、内容审核及多模态模型部署的研究者与工程师。

多模态基础模型的快速发展推动了跨模态理解与生成技术的进步,涵盖文本、图像、音频和视频等多种模态。然而,这些模型仍易受越狱攻击影响,可能绕过内置安全机制并生成有害内容。因此,深入理解越狱攻击手段及现有防御策略对保障多模态生成模型在真实场景中的安全部署至关重要,尤其在安全敏感应用中。本综述系统梳理了多模态生成模型中的越狱攻击与防御方法。首先,基于多模态越狱的通用生命周期,我们从输入、编码器、生成器和输出四个层级系统分析攻击与防御策略;在此基础上,构建了针对多模态生成模型的攻击方法、防御机制及评估框架的详细分类体系。此外,涵盖多种输入-输出配置,包括任意模态到文本、任意模态到视觉以及任意模态到任意模态的生成任务。最后,指出现有研究挑战,并提出未来研究方向。相关开源资源见:https://github.com/liuxuannan/Awesome-Multimodal-Jailbreak。

原文摘要 · Abstract (English)

The rapid evolution of multimodal foundation models has led to significant advancements in cross-modal understanding and generation across diverse modalities, including text, images, audio, and video. However, these models remain susceptible to jailbreak attacks, which can bypass built-in safety mechanisms and induce the production of potentially harmful content. Consequently, understanding the methods of jailbreak attacks and existing defense mechanisms is essential to ensure the safe deployment of multimodal generative models in real-world scenarios, particularly in security-sensitive applications. To provide comprehensive insight into this topic, this survey reviews jailbreak and defense in multimodal generative models. First, given the generalized lifecycle of multimodal jailbreak, we systematically explore attacks and corresponding defense strategies across four levels: input, encoder, generator, and output. Based on this analysis, we present a detailed taxonomy of attack methods, defense mechanisms, and evaluation frameworks specific to multimodal generative models. Additionally, we cover a wide range of input-output configurations, including modalities such as Any-to-Text, Any-to-Vision, and Any-to-Any within generative systems. Finally, we highlight current research challenges and propose potential directions for future research. The open-source repository corresponding to this work can be found at https://github.com/liuxuannan/Awesome-Multimodal-Jailbreak.

AI安全越狱攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。