用大模型自动生成攻击图文对,高效突破视觉语言模型安全防线
IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves

- 让大模型自己生成恶意图文组合进行黑盒攻击
- 对MiniGPT-4攻击成功率94%,仅需平均5.34次查询
- 构建首个多模态越狱基准,揭示主流模型安全短板
随着大型视觉语言模型(VLMs)日益重要,确保其安全部署成为关键。现有研究虽探索了对抗越狱攻击的鲁棒性,但受限于多模态数据多样性不足,普遍依赖从有害文本数据集中人工构造或对抗生成的图像,往往缺乏跨场景的有效性和多样性。本文提出IDEATOR,一种可自主生成恶意图文对的新型越狱方法。其核心思想是利用VLM自身作为红队模型,生成针对性越狱文本,并与由先进扩散模型生成的越狱图像配对。大量实验表明,IDEATOR在黑盒越狱中表现出高效率和强迁移能力:在MiniGPT-4上达到94%攻击成功率(ASR),平均仅需5.34次查询;在LLaVA、InstructBLIP和Chameleon上的转移攻击成功率分别为82%、88%和75%。基于IDEATOR的强迁移性与自动化流程,我们构建了VLJailbreakBench,一个包含3,654个多模态越狱样本的安全基准。11个最新发布VLMs的测试结果显示显著安全差距:挑战集在GPT-4o上达成46.31%的ASR,Claude-3.5-Sonnet为19.65%,凸显强化防御的紧迫性。该基准已公开于https://roywang021.github.io/VLJailbreakBench。
原文摘要 · Abstract (English)
As large Vision-Language Models (VLMs) gain prominence, ensuring their safe deployment has become critical. Recent studies have explored VLM robustness against jailbreak attacks-techniques that exploit model vulnerabilities to elicit harmful outputs. However, the limited availability of diverse multimodal data has constrained current approaches to rely heavily on adversarial or manually crafted images derived from harmful text datasets, which often lack effectiveness and diversity across different contexts. In this paper, we propose IDEATOR, a novel jailbreak method that autonomously generates malicious image-text pairs for black-box jailbreak attacks. IDEATOR is grounded in the insight that VLMs themselves could serve as powerful red team models for generating multimodal jailbreak prompts. Specifically, IDEATOR leverages a VLM to create targeted jailbreak texts and pairs them with jailbreak images generated by a state-of-the-art diffusion model. Extensive experiments demonstrate IDEATOR's high effectiveness and transferability, achieving a 94% attack success rate (ASR) in jailbreaking MiniGPT-4 with an average of only 5.34 queries, and high ASRs of 82%, 88%, and 75% when transferred to LLaVA, InstructBLIP, and Chameleon, respectively. Building on IDEATOR's strong transferability and automated process, we introduce the VLJailbreakBench, a safety benchmark comprising 3,654 multimodal jailbreak samples. Our benchmark results on 11 recently released VLMs reveal significant gaps in safety alignment. For instance, our challenge set achieves ASRs of 46.31% on GPT-4o and 19.65% on Claude-3.5-Sonnet, underscoring the urgent need for stronger defenses. VLJailbreakBench is publicly available at https://roywang021.github.io/VLJailbreakBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。