提出平衡主题相关性与异常强度的越狱方法,有效突破多模态模型安全防线。
Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- 通过重构恶意提示为语义关联的子任务,引入微妙异常信号
- 在13个模型上提升67%攻击成功率、21%有害输出率
- 揭示当前多模态安全机制的隐藏漏洞,适合安全研究者参考
多模态大语言模型(MLLM)广泛应用于视觉-语言推理任务,但其对对抗性提示的脆弱性仍是严重问题,安全机制常无法阻止有害输出。尽管近期越狱策略报告高成功率,但许多被判定为“成功”的响应实则无害、模糊或无关。这表明现有评估标准可能夸大了攻击效果。为此,我们提出四轴评估框架:输入主题相关性、输入分布外(OOD)强度、输出危害性与输出拒绝率,以识别真正有效的越狱。实证研究发现结构性权衡:高度相关的提示常被安全过滤器拦截,而过于偏离的提示虽可逃逸检测却难生成有害内容。但平衡相关性与新颖性的提示更易绕过检测并触发危险输出。基于此,我们开发递归重写策略Balanced Structural Decomposition(BSD),将恶意提示分解为语义对齐的子任务,同时注入细微的OOD信号和视觉线索,使输入更难被识别。在13个商用与开源MLLM上的测试显示,BSD持续提升攻击成功率、有害输出量并减少拒绝,相比此前方法,成功率提升67%,有害性提升21%,揭示当前多模态安全系统中未被充分重视的薄弱环节。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are widely used in vision-language reasoning tasks. However, their vulnerability to adversarial prompts remains a serious concern, as safety mechanisms often fail to prevent the generation of harmful outputs. Although recent jailbreak strategies report high success rates, many responses classified as "successful" are actually benign, vague, or unrelated to the intended malicious goal. This mismatch suggests that current evaluation standards may overestimate the effectiveness of such attacks. To address this issue, we introduce a four-axis evaluation framework that considers input on-topicness, input out-of-distribution (OOD) intensity, output harmfulness, and output refusal rate. This framework identifies truly effective jailbreaks. In a substantial empirical study, we reveal a structural trade-off: highly on-topic prompts are frequently blocked by safety filters, whereas those that are too OOD often evade detection but fail to produce harmful content. However, prompts that balance relevance and novelty are more likely to evade filters and trigger dangerous output. Building on this insight, we develop a recursive rewriting strategy called Balanced Structural Decomposition (BSD). The approach restructures malicious prompts into semantically aligned sub-tasks, while introducing subtle OOD signals and visual cues that make the inputs harder to detect. BSD was tested across 13 commercial and open-source MLLMs, where it consistently led to higher attack success rates, more harmful outputs, and fewer refusals. Compared to previous methods, it improves success rates by $67\%$ and harmfulness by $21\%$, revealing a previously underappreciated weakness in current multimodal safety systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。