用一张图骗VLM执行恶意指令,暴露视觉语言模型安全漏洞
VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands
- 用含恶意指令的对抗图像诱导模型响应
- 多个主流VLM在测试中均被成功绕过
- 小量毒内容可引发严重越界输出,适合安全研究者参考
视觉语言模型(VLMs)在多模态内容理解与生成方面表现突出,但其安全性仍面临严峻挑战。与纯文本模型不同,VLMs 因引入图像等多模态输入,产生新型漏洞如图像劫持,可诱使模型生成不当或有害内容。受文本类越狱攻击(如“立即做任何事”DAN)启发,本文提出VisualDAN:通过在单张对抗图像中嵌入类似DAN风格的恶意指令,先以肯定前缀(如“当然,我可以提供你需要的指导”)包装有害语料,再将该图像训练转化为文本域,从而诱发模型产生恶意输出。在MiniGPT-4、MiniGPT-v2、InstructBLIP和LLaVA等模型上的大量实验表明,VisualDAN能有效突破对齐后VLM的安全防护,迫使模型执行多种严重违反伦理标准的指令。结果还显示,一旦防御被攻破,极少量有毒内容即可显著放大有害输出。研究揭示了图像驱动攻击的紧迫性,为未来VLM对齐与安全研究提供了关键洞见。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have garnered significant attention for their remarkable ability to interpret and generate multimodal content. However, securing these models against jailbreak attacks continues to be a substantial challenge. Unlike text-only models, VLMs integrate additional modalities, introducing novel vulnerabilities such as image hijacking, which can manipulate the model into producing inappropriate or harmful responses. Drawing inspiration from text-based jailbreaks like the "Do Anything Now" (DAN) command, this work introduces VisualDAN, a single adversarial image embedded with DAN-style commands. Specifically, we prepend harmful corpora with affirmative prefixes (e.g., "Sure, I can provide the guidance you need") to trick the model into responding positively to malicious queries. The adversarial image is then trained on these DAN-inspired harmful texts and transformed into the text domain to elicit malicious outputs. Extensive experiments on models such as MiniGPT-4, MiniGPT-v2, InstructBLIP, and LLaVA reveal that VisualDAN effectively bypasses the safeguards of aligned VLMs, forcing them to execute a broad range of harmful instructions that severely violate ethical standards. Our results further demonstrate that even a small amount of toxic content can significantly amplify harmful outputs once the model's defenses are compromised. These findings highlight the urgent need for robust defenses against image-based attacks and offer critical insights for future research into the alignment and security of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。