用无关图像干扰可大幅降低编码攻击成功率,揭示防御机制新漏洞。
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
- 通过附加无意义图像,使文本编码攻击失效
- 最高降低73个百分点的攻击成功率,效果显著
- 适合研究模型安全机制与对抗样本的学者
我们发现视觉-语言模型(VLMs)中图像输入与现有黑盒防御之间存在反直觉交互:将编码的越狱提示与无关的干扰图像配对,可显著降低攻击成功率(ASR)。该效应源于防御流程变化,而非图像内容本身。在五个前沿VLMs、两种编码攻击家族和三种黑盒防御下,基于图像描述的防御(ECSO)在纯文本输入上几乎不改变ASR,但一旦附加无意义干扰图像,其ASR下降高达73个百分点;所有非饱和对比均在精确McNemar检验下显著。我们提出两个假设:防御分支依赖于图像存在性,且内在图像安全机制响应图像内容。三个控制实验排除了其他解释:空白画布与自然照片均有效,说明是图像存在而非内容所致;在无过滤层的开源VLM上仍有效,排除厂商过滤干扰;使用非符号化语义编码器也复现该现象,表明不特异于符号混淆。但无条件附加干扰会令良性拒绝率升至20%–79%,增加10–67个百分点。若仅由轻量级编码输入检测器触发附加,则可维持原始拒绝率并保留安全收益,此时检测器召回率成为关键约束。在针对描述重检的自适应攻击下,效果减弱但仍保持有效。本研究定位为对防御流程交互的观察,而非可靠防御方案。
原文摘要 · Abstract (English)
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). The operative change is in the defense pipeline, not in the image. Across five frontier VLMs, two encoded-attack families, and three black-box defenses, a caption-mediated defense (ECSO) that leaves ASR essentially unchanged on text-only encoded input drops it by up to $73$pp once a content-free decoy is attached; every non-saturated contrast is significant under exact McNemar tests. We advance two hypotheses for this pattern, supported by indirect evidence rather than pipeline introspection, since a black-box threat model precludes inspecting vendor internals: caption-mediated defenses branch on image presence, and intrinsic image-side safety engages on image-resident content. Three controls constrain the explanation. Blank-canvas and natural-photograph decoys reproduce the effect on every model, implicating image presence rather than content; the effect replicates on three open-weight VLMs served with no moderation layer, so it is not a vendor-filtering artifact; and a non-symbolic, meaning-based encoder reproduces it, so it is not specific to symbolic obfuscation. Attaching a decoy unconditionally is not deployable --- it raises benign refusal to $20$--$79\%$, an inflation of $+10$ to $+67$pp --- but gating attachment on a lightweight encoded-input detector returns benign refusal to the text baseline while preserving the safety gain wherever the detector fires, making detector recall the binding constraint. Under adaptive attacks that target the caption-mediated re-check, the effect degrades but holds. We frame this as an observation about pipeline interaction, not as a robust defense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。