拆解多模态越狱漏洞的四大因素,精准定位安全弱点。
MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities

- 分离危害意图、提示框架、视觉语义和指令载体,系统化测试各因素影响。
- 提示框架是主要变量,权威类视觉线索显著提升越狱成功率。
- 适合安全研究者与模型开发者用于评估和优化多模态系统鲁棒性。
多模态大语言模型(MLLMs)在实际应用中日益普及,但不同因素如何影响其越狱漏洞仍不明确。现有基准将有害意图、提示框架、视觉语义和指令载体混合在同一实例中,掩盖了具体漏洞来源。为此,我们提出MMJailBench,一个因子解耦的基准,通过受控配置系统性地变化与组合这些因素,实现细粒度比较与因子级归因。对16个开源与专有MLLM的大规模评估显示,漏洞表现高度异质且依赖模型。不同危害领域间漏洞差异显著,现有对齐策略覆盖不均。提示框架是主要变异源,任务相关的视觉语义系统性提高越狱风险,具有权威特征的视觉线索暴露更明显脆弱性,而视觉呈现指令并未比直接文本指令更易引发越狱。为进一步探究多模态上下文带来的风险,我们对代表性开源模型进行诊断分析,识别出内部表征与跨模态交互中的脆弱模式。最后,我们开发了一个模块化多模态越狱评估套件,支持完整与轻量配置、多种评判方式及多维指标,实现可复现、可扩展、低成本的多模态越狱审计。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple harmful intent, prompt framing, visual semantics, and instruction carrier within individual jailbreak instances, obscuring the specific sources of observed vulnerabilities. To address this limitation, we introduce MMJailBench, a factorized benchmark that systematically varies and combines these factors under controlled configurations, enabling fine-grained comparison and factor-level attribution. Large-scale evaluations across 16 open-weight and proprietary MLLMs reveal highly heterogeneous and model-dependent vulnerability profiles. Jailbreak vulnerability varies markedly across harm domains, exposing uneven coverage in current multimodal safety alignment. Prompt framing emerges as the dominant source of variation, task-relevant visual semantics systematically increase jailbreak susceptibility with authority-like cues exposing particularly pronounced vulnerabilities, and visually rendered instructions do not consistently increase jailbreak susceptibility relative to direct textual instructions. To further investigate the risks introduced by multimodal context, we conduct diagnostic analyses on a representative open-weight model and identify vulnerability-associated patterns in internal representations and cross-modal interactions. Finally, we develop a modular multimodal jailbreak evaluation suite with full and lightweight configurations, multiple judge options, and multidimensional metrics, enabling reproducible, scalable, and cost-efficient multimodal jailbreak auditing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。