利用多模态推理绕过视觉语言模型安全限制
Jailbreaks on Vision Language Model via Multimodal Reasoning
- 通过链式思维提示构建隐蔽攻击文本
- 自适应噪声扰动使攻击成功率提升,图像仍自然
- 适合研究安全漏洞与对抗攻击的学者参考
视觉语言模型(VLMs)在视觉问答、图像描述和文生图等任务中发挥核心作用。然而,其输出对提示变化极为敏感,暴露出安全对齐的潜在漏洞。本文提出一种劫持框架,利用后训练链式思维(CoT)提示构造可规避安全过滤的隐蔽提示。为进一步提升攻击成功率(ASR),提出基于ReAct的自适应噪声机制,根据模型反馈迭代扰动输入图像,聚焦于最可能触发安全防御的区域,从而增强隐蔽性与逃逸能力。实验表明,该双策略显著提升攻击成功率,同时保持文本与视觉内容的自然性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have become central to tasks such as visual question answering, image captioning, and text-to-image generation. However, their outputs are highly sensitive to prompt variations, which can reveal vulnerabilities in safety alignment. In this work, we present a jailbreak framework that exploits post-training Chain-of-Thought (CoT) prompting to construct stealthy prompts capable of bypassing safety filters. To further increase attack success rates (ASR), we propose a ReAct-driven adaptive noising mechanism that iteratively perturbs input images based on model feedback. This approach leverages the ReAct paradigm to refine adversarial noise in regions most likely to activate safety defenses, thereby enhancing stealth and evasion. Experimental results demonstrate that the proposed dual-strategy significantly improves ASR while maintaining naturalness in both text and visual domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。