arXiv:2601.05339cs.CRcs.AI2026-01被引 2

提出多轮越狱攻击与防御框架,提升多模态大模型安全性

Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models

  • 设计多轮提示越狱攻击,挖掘模型漏洞
  • 提出FragGuard防御机制,有效抵御越狱攻击
  • 在多个主流模型上验证攻防效果,适合安全研究者

近年来,多模态大语言模型(MLLMs)的安全漏洞成为生成式人工智能研究中的严重问题。这些具备高精度多模态任务能力的智能模型,也极易受到精心设计的安全攻击,如越狱攻击,导致模型行为被操纵并绕过安全约束。本文提出MJAD-MLLMs框架,系统分析多轮越狱攻击及基于多大语言模型的防御技术。主要贡献包括:第一,提出一种新型多轮越狱攻击,针对多轮提示下的MLLMs漏洞;第二,提出一种片段优化的多大语言模型防御机制FragGuard,有效缓解越狱攻击;第三,在多个前沿开源与闭源MLLMs及基准数据集上进行大量实验,评估所提攻防方案效果,并与现有技术对比。

原文摘要 · Abstract (English)

In recent years, the security vulnerabilities of Multi-modal Large Language Models (MLLMs) have become a serious concern in the Generative Artificial Intelligence (GenAI) research. These highly intelligent models, capable of performing multi-modal tasks with high accuracy, are also severely susceptible to carefully launched security attacks, such as jailbreaking attacks, which can manipulate model behavior and bypass safety constraints. This paper introduces MJAD-MLLMs, a holistic framework that systematically analyzes the proposed Multi-turn Jailbreaking Attacks and multi-LLM-based defense techniques for MLLMs. In this paper, we make three original contributions. First, we introduce a novel multi-turn jailbreaking attack to exploit the vulnerabilities of the MLLMs under multi-turn prompting. Second, we propose a novel fragment-optimized and multi-LLM defense mechanism, called FragGuard, to effectively mitigate jailbreaking attacks in the MLLMs. Third, we evaluate the efficacy of the proposed attacks and defenses through extensive experiments on several state-of-the-art (SOTA) open-source and closed-source MLLMs and benchmark datasets, and compare their performance with the existing techniques.

越狱攻击多模态模型安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。