用语义无关输入突破多模态大模型的图像安全边界
Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs
- 通过重构-生成策略,将恶意意图隐藏在中性输入中
- 对GPT-5实现98.21%的攻击成功率,验证视觉安全缺陷
- 揭示当前多模态模型在图像生成安全上的重大漏洞
多模态大语言模型(MLLMs)的快速发展带来了复杂的安全部署挑战,尤其在文本与图像安全的交界处。现有研究虽已探索了MLLMs的安全漏洞,但对其视觉安全边界的探究仍显不足。本文提出「超越视觉安全」(BVS)框架,一种专门用于探测MLLMs视觉安全边界的图像-文本对越狱方法。BVS采用“重构-生成”策略,结合中性化视觉拼接与归纳重组技术,将恶意意图从原始输入中剥离,从而诱导模型生成有害图像。实验结果表明,BVS对GPT-5(2026年1月12日发布版)实现了98.21%的越狱成功率。研究揭示了当前MLLMs在视觉安全对齐方面存在的关键漏洞。
原文摘要 · Abstract (English)
The rapid advancement of Multimodal Large Language Models (MLLMs) has introduced complex security challenges, particularly at the intersection of textual and visual safety. While existing schemes have explored the security vulnerabilities of MLLMs, the investigation into their visual safety boundaries remains insufficient. In this paper, we propose Beyond Visual Safety (BVS), a novel image-text pair jailbreaking framework specifically designed to probe the visual safety boundaries of MLLMs. BVS employs a "reconstruction-then-generation" strategy, leveraging neutralized visual splicing and inductive recomposition to decouple malicious intent from raw inputs, thereby leading MLLMs to be induced into generating harmful images. Experimental results demonstrate that BVS achieves a remarkable jailbreak success rate of 98.21\% against GPT-5 (12 January 2026 release). Our findings expose critical vulnerabilities in the visual safety alignment of current MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。