arXiv:2508.05087cs.MMcs.AI2025-08被引 11

通过图像扰动与文本引导协同攻击,让多模态大模型生成真正有害的内容。

JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering

  • 用视觉扰动+文本引导联合突破安全防护
  • 在多个模型上实现98%以上攻击成功率和85%恶意意图达成率
  • 适合研究安全漏洞或对抗攻击的人员参考

针对多模态大语言模型(MLLMs)的越狱攻击是当前重要研究方向。现有方法多关注攻击成功率(ASR),却常忽略生成内容是否真正满足攻击者恶意意图,导致输出质量差、仅能绕过过滤但无实质危害。为此,我们提出JPS(Jailbreak MLLMs with collaborative visual Perturbation and textual Steering),通过视觉图像扰动与文本引导提示的协同作用实现越狱。JPS采用目标导向的对抗性图像扰动以有效绕过安全机制,并利用多智能体系统优化“引导提示”,精准控制大模型生成符合攻击意图的内容。视觉与文本组件通过迭代联合优化提升性能。为评估攻击输出质量,我们提出恶意意图达成率(MIFR)指标,由基于推理的LLM评估器衡量。实验表明,JPS在多种MLLM和基准测试中均达到新的最优水平,同时在攻击成功率和恶意意图达成率上显著优于现有方法。代码已公开于https://github.com/thu-coai/JPS。注意:本文含潜在敏感内容。

原文摘要 · Abstract (English)

Jailbreak attacks against multimodal large language Models (MLLMs) are a significant research focus. Current research predominantly focuses on maximizing attack success rate (ASR), often overlooking whether the generated responses actually fulfill the attacker's malicious intent. This oversight frequently leads to low-quality outputs that bypass safety filters but lack substantial harmful content. To address this gap, we propose JPS, \underline{J}ailbreak MLLMs with collaborative visual \underline{P}erturbation and textual \underline{S}teering, which achieves jailbreaks via corporation of visual image and textually steering prompt. Specifically, JPS utilizes target-guided adversarial image perturbations for effective safety bypass, complemented by "steering prompt" optimized via a multi-agent system to specifically guide LLM responses fulfilling the attackers' intent. These visual and textual components undergo iterative co-optimization for enhanced performance. To evaluate the quality of attack outcomes, we propose the Malicious Intent Fulfillment Rate (MIFR) metric, assessed using a Reasoning-LLM-based evaluator. Our experiments show JPS sets a new state-of-the-art in both ASR and MIFR across various MLLMs and benchmarks, with analyses confirming its efficacy. Codes are available at \href{https://github.com/thu-coai/JPS}{https://github.com/thu-coai/JPS}. \color{warningcolor}{Warning: This paper contains potentially sensitive contents.}

越狱攻击多模态对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。