arXiv:2512.20168cs.CRcs.AI2025-12中稿 · Network and Distri…被引 8

用双隐写术隐藏恶意指令,突破商业多模态大模型安全防线

Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography

  • 通过图像中嵌入隐蔽指令和响应,实现跨模态攻击
  • 在真实系统上实现最高99%的越狱成功率
  • 揭示现有防御对多模态攻击的盲区,适合安全研究者参考

将语言理解与视觉等感知模态结合的多模态大语言模型(MLLM)是现代AI系统的核心,尤其在开放交互环境中运行的智能体。然而,其普及也带来滥用风险,如生成有害内容。尽管通过对齐技术使模型行为符合人类价值观,但近期研究显示越狱攻击仍可绕过这些机制。当前多数越狱方法针对开源模型,对部署了额外过滤器的商业MLLM集成系统效果有限。这些过滤器依赖一个关键假设:恶意内容必须在输入或输出中显式存在。这一假设在传统LLM系统中成立,但在多模态系统中失效——攻击者可利用多模态特性隐藏恶意意图,造成虚假安全。为此,我们提出Odysseus,一种新型越狱范式,引入双隐写术,将恶意查询与响应隐含于看似无害的图像中。在基准数据集上的实验表明,Odysseus成功越狱多个前沿且真实的商业级MLLM集成系统,攻击成功率高达99%。该工作暴露了现有防御的根本性盲点,呼吁重新思考多模态系统中的跨模态安全问题。

原文摘要 · Abstract (English)

By integrating language understanding with perceptual modalities such as images, multimodal large language models (MLLMs) constitute a critical substrate for modern AI systems, particularly intelligent agents operating in open and interactive environments. However, their increasing accessibility also raises heightened risks of misuse, such as generating harmful or unsafe content. To mitigate these risks, alignment techniques are commonly applied to align model behavior with human values. Despite these efforts, recent studies have shown that jailbreak attacks can circumvent alignment and elicit unsafe outputs. Currently, most existing jailbreak methods are tailored for open-source models and exhibit limited effectiveness against commercial MLLM-integrated systems, which often employ additional filters. These filters can detect and prevent malicious input and output content, significantly reducing jailbreak threats. In this paper, we reveal that the success of these safety filters heavily relies on a critical assumption that malicious content must be explicitly visible in either the input or the output. This assumption, while often valid for traditional LLM-integrated systems, breaks down in MLLM-integrated systems, where attackers can leverage multiple modalities to conceal adversarial intent, leading to a false sense of security in existing MLLM-integrated systems. To challenge this assumption, we propose Odysseus, a novel jailbreak paradigm that introduces dual steganography to covertly embed malicious queries and responses into benign-looking images. Extensive experiments on benchmark datasets demonstrate that our Odysseus successfully jailbreaks several pioneering and realistic MLLM-integrated systems, achieving up to 99% attack success rate. It exposes a fundamental blind spot in existing defenses, and calls for rethinking cross-modal security in MLLM-integrated systems.

越狱攻击多模态隐写术安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。