arXiv:2603.17372cs.CVcs.AI2026-03被引 3

图像会诱导视觉语言模型进入越狱状态,导致安全失效。

Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift

  • 发现越狱行为源于视觉输入引发的表征偏移。
  • 提出越狱相关偏移量(JRS),可精准量化越狱程度。
  • 在不损害正常任务性能前提下,有效防御多种越狱攻击。

大型视觉语言模型(VLMs)在引入视觉模态后,安全对齐能力常被削弱。即使文本提示明确包含恶意意图,加入图像后仍会显著提高越狱成功率。本文观察到,VLMs 在表征空间中能清晰区分良性输入与有害输入;且在有害输入中,越狱样本形成独立于拒绝样本的内部状态。这表明越狱并非因无法识别恶意意图,而是视觉模态将表征推向特定越狱状态,从而触发失败拒绝。为量化此过程,我们识别出越狱方向,并定义图像诱导的表征偏移中沿该方向的分量为越狱相关偏移(JRS)。分析显示,JRS 能可靠刻画越狱行为,统一解释多种越狱场景。最后,我们提出一种推理时移除JRS的防御方法(JRS-Rem),实验表明其在多种场景下均具强防御效果,同时保持良性任务性能。

原文摘要 · Abstract (English)

Large vision-language models (VLMs) often exhibit weakened safety alignment with the integration of the visual modality. Even when text prompts contain explicit harmful intent, adding an image can substantially increase jailbreak success rates. In this paper, we observe that VLMs can clearly distinguish benign inputs from harmful ones in their representation space. Moreover, even among harmful inputs, jailbreak samples form a distinct internal state that is separable from refusal samples. These observations suggest that jailbreaks do not arise from a failure to recognize harmful intent. Instead, the visual modality shifts representations toward a specific jailbreak state, thereby leading to a failure to trigger refusal. To quantify this transition, we identify a jailbreak direction and define the jailbreak-related shift as the component of the image-induced representation shift along this direction. Our analysis shows that the jailbreak-related shift reliably characterizes jailbreak behavior, providing a unified explanation for diverse jailbreak scenarios. Finally, we propose a defense method that enhances VLM safety by removing the jailbreak-related shift (JRS-Rem) at inference time. Experiments show that JRS-Rem provides strong defense across multiple scenarios while preserving performance on benign tasks.

视觉语言模型越狱攻击安全对齐表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。