用视觉方式绕过AI安全限制,实验证明图像比文本更易被攻击
Jailbreaking Vision-Language Models Through the Visual Modality

- 通过图像符号、替换物、视觉隐喻等方法诱导模型生成有害内容
- 在6个前沿模型上成功率达40.9%,远超同等文本攻击的10.7%
- 揭示图文对齐缺陷,提醒安全训练需重视视觉模态
视觉语言模型(VLMs)的视觉模态是未被充分探索的安全漏洞。我们提出四种利用视觉成分的越狱攻击:(1) 将有害指令编码为视觉符号序列并附解码说明;(2) 用无害物体(如炸弹→香蕉)替代有害对象,再用替代词触发有害行为;(3) 替换图像中(如书封)的有害文字为无害词汇,但保留视觉语境含义;(4) 设计视觉类比谜题,解题需推断被禁止概念。在六个前沿VLM上评估显示,这些视觉攻击可绕过安全对齐,暴露跨模态对齐缺口:基于文本的安全训练无法自动泛化至视觉传达的有害意图。例如,我们的视觉密码在Claude-Haiku-4.5上成功率达40.9%,而等效文本密码仅10.7%。为进一步理解攻击机制,我们提供初步可解释性与缓解结果。研究强调,鲁棒VLM对齐必须将视觉视为安全后训练的核心目标。
原文摘要 · Abstract (English)
The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak attacks exploiting the vision component: (1) encoding harmful instructions as visual symbol sequences with a decoding legend, (2) replacing harmful objects with benign substitutes (e.g., bomb -> banana) then prompting for harmful actions using the substitute term, (3) replacing harmful text in images (e.g., on book covers) with benign words while visual context preserves the original meaning, and (4) visual analogy puzzles whose solution requires inferring a prohibited concept. Evaluating across six frontier VLMs, our visual attacks bypass safety alignment and expose a cross-modality alignment gap: text-based safety training does not automatically generalize to harmful intent conveyed visually. For example, our visual cipher achieves 40.9% attack success on Claude-Haiku-4.5 versus 10.7% for an equivalent textual cipher. To further our insight into the attack mechanism, we present preliminary interpretability and mitigation results. These findings highlight that robust VLM alignment requires treating vision as a first-class target for safety post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。