arXiv:2504.01444cs.CRcs.AI2025-04中稿 · IEEE International…被引 10

用图像代码诱导漏洞,突破多模态大模型安全防线

PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization

  • 通过图像中的代码样式指令隐藏恶意意图,逐层绕过防御
  • 在Gemini-Pro Vision上实现84.13%攻击成功率,高于此前方法
  • 揭示当前多模态模型安全短板,适合安全研究者参考

多模态大语言模型(MLLM)融合视觉等模态,显著提升AI能力,但也引入新安全风险。利用视觉模态漏洞和代码训练数据的长尾分布特性,本文提出PiCo框架,通过分层攻击策略逐步突破高级MLLM的多级防御。该方法采用词元级排版攻击规避输入过滤,并将有害意图嵌入编程上下文指令以绕过运行时监控。为全面评估攻击影响,提出新评估指标,同时衡量输出毒性与有用性。通过在代码风格的视觉指令中嵌入恶意意图,PiCo在Gemini-Pro Vision上实现84.13%的平均攻击成功率,在GPT-4上达52.66%,超越已有方法。实验凸显现有防御体系的关键缺陷,强调需构建更稳健的安全机制来保护先进MLLM。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs), which integrate vision and other modalities into Large Language Models (LLMs), significantly enhance AI capabilities but also introduce new security vulnerabilities. By exploiting the vulnerabilities of the visual modality and the long-tail distribution characteristic of code training data, we present PiCo, a novel jailbreaking framework designed to progressively bypass multi-tiered defense mechanisms in advanced MLLMs. PiCo employs a tier-by-tier jailbreak strategy, using token-level typographic attacks to evade input filtering and embedding harmful intent within programming context instructions to bypass runtime monitoring. To comprehensively assess the impact of attacks, a new evaluation metric is further proposed to assess both the toxicity and helpfulness of model outputs post-attack. By embedding harmful intent within code-style visual instructions, PiCo achieves an average Attack Success Rate (ASR) of 84.13% on Gemini-Pro Vision and 52.66% on GPT-4, surpassing previous methods. Experimental results highlight the critical gaps in current defenses, underscoring the need for more robust strategies to secure advanced MLLMs.

多模态安全模型攻击视觉编码对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。