显式调用图像工具能显著提升多模态越狱攻击的防御能力。
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

- 通过显式调用图像工具,模型更难被越狱攻击成功。
- 相比直接回答,攻击成功率平均降低约30%。
- 适合关注多模态安全与模型防护的研究者。
思考式图像推理正成为大视觉语言模型的新推理范式,但其安全影响仍不明确。现有系统采用多种流程设计,包括直接生成回答、仅文本前导、视觉状态操作以及显式外部图像工具调用。本文探究哪种范式能提升多模态越狱攻击的鲁棒性。实验表明,在多个视觉语言模型上,显式图像工具交互使攻击成功率最低,平均相对降低约30%。这一结果出人意料:即使返回的图像工具输出被手动覆盖或看似不安全,攻击成功率依然较低;但在仅文本前导控制下,成功率接近直接回答水平。这表明低攻击成功率并非源于良性图像语义或文本工具痕迹本身。为此,我们提出图像工具安全向量框架,将图像工具调用建模为隐藏表示向安全相关方向的残差偏移。表示层面分析与激活干预验证了该解释。总体而言,显式图像工具交互是提升越狱鲁棒性的有前景设计模式,同时呼吁进行管道特定的安全评估。
原文摘要 · Abstract (English)
Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response generation, text-only prior turn, visual-state manipulation, and explicit external image-tool invocation. In this paper, we ask which of these evaluated paradigms improves multimodal jailbreak robustness, and why. Across multiple vision-language models, explicit image-tool interaction yields the lowest attack success rates in our experiments, reducing jailbreak success by around 30% relative on average across the evaluated models. This finding is initially surprising: ASR remains low even when the returned image-tool output is manually overridden or itself unsafe-looking, but returns near direct-answering levels under text-only prior turn controls. These results indicate that the lower ASR is not explained by benign returned-image semantics or by the textual image-tool trace alone. To explain the pattern, we introduce an image-tool safety vector framework that models image-tool invocation as a residual shift in hidden representations toward a safety-relevant direction. Representation-level analyses and activation interventions support this account. Overall, our results suggest that explicit image-tool interaction is a promising design pattern for improving jailbreak robustness, while also motivating pipeline-specific safety evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。