用图像隐写术在视觉语言模型中悄悄植入恶意指令,成功率超90%。
Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
- 通过图像最低有效位隐写技术,将恶意指令嵌入图片中
- 在GPT-4o和Gemini-1.5 Pro上仅用平均3次查询,成功率超90%
- 适配多种商业模型,适合研究安全漏洞与对抗攻击的学者
多模态大语言模型(MLLMs)具备强大的跨模态推理能力,但更广的输入空间也带来了新的攻击面。以往的越狱攻击常通过文本向视觉等低对齐模态注入恶意指令,但随着模型引入更多跨模态一致性机制,这类显式攻击易被检测和拦截。本文提出一种新型隐式越狱框架IJA,通过最低有效位隐写术将恶意指令隐蔽嵌入图像,并搭配看似无害的图像相关文本提示。为提升对不同MLLM的攻击效果,引入由代理模型生成的对抗后缀,以及基于模型反馈迭代优化提示与嵌入的模板优化模块。在GPT-4o和Gemini-1.5 Pro等商用模型上,本方法平均仅需3次查询,攻击成功率超过90%。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) enable powerful cross-modal reasoning capabilities. However, the expanded input space introduces new attack surfaces. Previous jailbreak attacks often inject malicious instructions from text into less aligned modalities, such as vision. As MLLMs increasingly incorporate cross-modal consistency and alignment mechanisms, such explicit attacks become easier to detect and block. In this work, we propose a novel implicit jailbreak framework termed IJA that stealthily embeds malicious instructions into images via least significant bit steganography and couples them with seemingly benign, image-related textual prompts. To further enhance attack effectiveness across diverse MLLMs, we incorporate adversarial suffixes generated by a surrogate model and introduce a template optimization module that iteratively refines both the prompt and embedding based on model feedback. On commercial models like GPT-4o and Gemini-1.5 Pro, our method achieves attack success rates of over 90% using an average of only 3 queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。