arXiv:2507.13761cs.CL2025-07

研究视觉语言模型的提示敏感性,发现三类设计可轻易触发有害输出。

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models

  • 通过分析提示中的视觉信息、对抗样本和积极开头,揭示多模态漏洞。
  • 仅需三个上下文示例即可成功越狱,且使用普通图像也能生效。
  • 提出跳接结构增强攻击效果,揭示表情包也具高风险性。

语言模型对提示设计极为敏感——输入微小变化可导致输出大幅偏离。本文探究提示设计中离散组件如何影响视觉语言模型(VLM)生成不当内容的能力。具体分析三个关键因素:(a) 详细视觉信息的包含,(b) 对抗样本的存在,以及 (c) 积极语气开头的使用。结果表明,尽管在单模态场景(仅文本或仅图像)下,VLM 能可靠区分良性与有害输入,但在多模态环境下这一能力显著下降。上述三因素独立即可触发越狱,且仅需三个上下文示例(as few as three)即可促使模型生成不当内容。此外,我们提出一种框架,利用VLM内部两层之间的跳接连接,大幅提升越狱成功率,即使使用良性图像亦可实现。最后,我们证明,常被视为幽默或无害的表情包,其诱发有害内容的效果可与有毒视觉内容相当,凸显了VLM潜在的隐蔽复杂漏洞。

原文摘要 · Abstract (English)

Language models are highly sensitive to prompt formulations - small changes in input can drastically alter their output. This raises a critical question: To what extent can prompt sensitivity be exploited to generate inapt content? In this paper, we investigate how discrete components of prompt design influence the generation of inappropriate content in Visual Language Models (VLMs). Specifically, we analyze the impact of three key factors on successful jailbreaks: (a) the inclusion of detailed visual information, (b) the presence of adversarial examples, and (c) the use of positively framed beginning phrases. Our findings reveal that while a VLM can reliably distinguish between benign and harmful inputs in unimodal settings (text-only or image-only), this ability significantly degrades in multimodal contexts. Each of the three factors is independently capable of triggering a jailbreak, and we show that even a small number of in-context examples (as few as three) can push the model toward generating inappropriate outputs. Furthermore, we propose a framework that utilizes a skip-connection between two internal layers of the VLM, which substantially increases jailbreak success rates, even when using benign images. Finally, we demonstrate that memes, often perceived as humorous or harmless, can be as effective as toxic visuals in eliciting harmful content, underscoring the subtle and complex vulnerabilities of VLMs.

视觉语言模型越狱攻击提示安全多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。