研究多模态模型的攻击方式及漏洞传播风险。
Attacks on multimodal models
- 聚焦CLIP图像编码器,测试多种补丁攻击效果。
- 发现开源组件漏洞可被继承并影响下游多模态模型。
- 适合关注AI安全与工业部署风险的研究者阅读。
当前支持多模态协同的对话式模型日益流行,但其潜在攻击风险不容忽视,尤其当模型包含开源组件时。本文研究此类模型面临的各类攻击,并评估其泛化能力。现代视觉语言模型(如LLaVA、BLIP)常复用其他模型的预训练部分,因此本研究重点聚焦于CLIP架构及其图像编码器(CLIP-ViT),以及针对该结构的多种补丁攻击变体。
原文摘要 · Abstract (English)
Today, models capable of working with various modalities simultaneously in a chat format are gaining increasing popularity. Despite this, there is an issue of potential attacks on these models, especially considering that many of them include open-source components. It is important to study whether the vulnerabilities of these components are inherited and how dangerous this can be when using such models in the industry. This work is dedicated to researching various types of attacks on such models and evaluating their generalization capabilities. Modern VLM models (LLaVA, BLIP, etc.) often use pre-trained parts from other models, so the main part of this research focuses on them, specifically on the CLIP architecture and its image encoder (CLIP-ViT) and various patch attack variations for it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。