用故事化角色沉浸突破多模态模型安全防护,成功率提升17.5%
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
- 通过环境-角色-动作三元组构建多轮视觉叙事引导攻击
- 在六款主流多模态大模型上攻击成功率最高提升17.5%
- 揭示角色代入与语义重构可激活模型隐性偏见
尽管安全机制在过滤有害文本输入方面已显著进步,但多模态大语言模型(MLLMs)仍易受利用跨模态推理能力的多模态越狱攻击。本文提出MIRAGE框架,通过叙事驱动的上下文与角色沉浸,绕过MLLM的安全机制。该方法将有害请求分解为环境、角色和动作三元组,借助Stable Diffusion生成图像与文本组成的多轮视觉叙事序列,引导目标模型进入引人入胜的侦探故事中。这一过程逐步降低模型防御,并通过结构化上下文线索微妙引导其推理,最终诱使产生有害响应。在选定数据集上对六款主流MLLM的广泛实验表明,MIRAGE达到当前最佳性能,攻击成功率较最优基线最高提升17.5%。此外,我们证明角色沉浸与结构化语义重构可激活模型内在偏见,促使模型自发违反伦理约束。这些结果揭示了当前多模态安全机制的关键漏洞,凸显应对跨模态威胁的紧迫性。
原文摘要 · Abstract (English)
While safety mechanisms have significantly progressed in filtering harmful text inputs, MLLMs remain vulnerable to multimodal jailbreaks that exploit their cross-modal reasoning capabilities. We present MIRAGE, a novel multimodal jailbreak framework that exploits narrative-driven context and role immersion to circumvent safety mechanisms in Multimodal Large Language Models (MLLMs). By systematically decomposing the toxic query into environment, role, and action triplets, MIRAGE constructs a multi-turn visual storytelling sequence of images and text using Stable Diffusion, guiding the target model through an engaging detective narrative. This process progressively lowers the model's defences and subtly guides its reasoning through structured contextual cues, ultimately eliciting harmful responses. In extensive experiments on the selected datasets with six mainstream MLLMs, MIRAGE achieves state-of-the-art performance, improving attack success rates by up to 17.5% over the best baselines. Moreover, we demonstrate that role immersion and structured semantic reconstruction can activate inherent model biases, facilitating the model's spontaneous violation of ethical safeguards. These results highlight critical weaknesses in current multimodal safety mechanisms and underscore the urgent need for more robust defences against cross-modal threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。