无需标签即可对视觉语言模型发动大规模攻击
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
- 基于LAION-400M数据自监督训练,无需目标标签即可生成攻击
- 在5个开源模型上成功攻击,且可跨模型迁移至商业系统
- 揭示视觉语言模型普遍存在系统性安全漏洞
由于多模态能力,视觉语言模型(VLMs)已在真实场景中广泛应用。然而,近期研究发现其易受基于图像的对抗攻击。传统定向攻击需特定目标和标签,限制了实际影响。本文提出AnyAttack,一种自监督框架,通过在无标签的LAION-400M数据集上预训练,实现前所未有的灵活性——任意图像均可转化为针对任意输出的攻击向量,适用于多种VLM。该方法从根本上改变威胁格局,使对抗能力以空前规模可得。我们在五个开源VLM(CLIP、BLIP、BLIP2、InstructBLIP、MiniGPT-4)上进行了广泛验证,证明AnyAttack在多样化多模态任务中的有效性。最令人担忧的是,该攻击可无缝迁移到Google Gemini、Claude Sonnet、Microsoft Copilot和OpenAI GPT等商业系统,暴露出亟需关注的系统性漏洞。
原文摘要 · Abstract (English)
Due to their multimodal capabilities, Vision-Language Models (VLMs) have found numerous impactful applications in real-world scenarios. However, recent studies have revealed that VLMs are vulnerable to image-based adversarial attacks. Traditional targeted adversarial attacks require specific targets and labels, limiting their real-world impact.We present AnyAttack, a self-supervised framework that transcends the limitations of conventional attacks through a novel foundation model approach. By pre-training on the massive LAION-400M dataset without label supervision, AnyAttack achieves unprecedented flexibility - enabling any image to be transformed into an attack vector targeting any desired output across different VLMs.This approach fundamentally changes the threat landscape, making adversarial capabilities accessible at an unprecedented scale. Our extensive validation across five open-source VLMs (CLIP, BLIP, BLIP2, InstructBLIP, and MiniGPT-4) demonstrates AnyAttack's effectiveness across diverse multimodal tasks. Most concerning, AnyAttack seamlessly transfers to commercial systems including Google Gemini, Claude Sonnet, Microsoft Copilot and OpenAI GPT, revealing a systemic vulnerability requiring immediate attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。