多模态模型的视觉输入易被伪装攻击,导致错误输出。
Seeing is Deceiving: Exploitation of Visual Pathways in Multi-Modal Language Models
- 分析了视觉输入被篡改的多种攻击方式
- 部分攻击在视觉上几乎与原图无异却能误导模型
- 适合关注AI安全与防御的研究者和开发者
多模态语言模型(MLLMs)通过融合视觉与文本数据,推动了图像描述、视觉问答及多模态内容生成等应用的发展,在医疗、自动驾驶和数字内容等领域具有重要价值。然而,多源数据融合也带来安全风险:攻击者可通过操纵视觉或文本输入,甚至两者联合,诱导模型产生非预期甚至有害响应。本文系统梳理了针对MLLM视觉路径的各类攻击策略,分为简单视觉扰动、跨模态干扰,以及高级攻击如VLATTACK、HADES和协同多模态对抗攻击(Co-Attack)。这些攻击在视觉上与原始图像极为相似,却可误导当前最鲁棒的模型。我们还讨论了隐私泄露与关键应用中的安全威胁。针对防御,综述了SmoothVLM、像素级随机化、MirrorCheck等方法的优劣,并提出自适应防御、更优评估工具及双模态保护机制等新方向。本综述旨在整合最新进展,推动构建更安全可靠的多模态AI系统。
原文摘要 · Abstract (English)
Multi-Modal Language Models (MLLMs) have transformed artificial intelligence by combining visual and text data, making applications like image captioning, visual question answering, and multi-modal content creation possible. This ability to understand and work with complex information has made MLLMs useful in areas such as healthcare, autonomous systems, and digital content. However, integrating multiple types of data also creates security risks. Attackers can manipulate either the visual or text inputs, or both, to make the model produce unintended or even harmful responses. This paper reviews how visual inputs in MLLMs can be exploited by various attack strategies. We break down these attacks into categories: simple visual tweaks and cross-modal manipulations, as well as advanced strategies like VLATTACK, HADES, and Collaborative Multimodal Adversarial Attack (Co-Attack). These attacks can mislead even the most robust models while looking nearly identical to the original visuals, making them hard to detect. We also discuss the broader security risks, including threats to privacy and safety in important applications. To counter these risks, we review current defense methods like the SmoothVLM framework, pixel-wise randomization, and MirrorCheck, looking at their strengths and limitations. We also discuss new methods to make MLLMs more secure, including adaptive defenses, better evaluation tools, and security approaches that protect both visual and text data. By bringing together recent developments and identifying key areas for improvement, this review aims to support the creation of more secure and reliable multi-modal AI systems for real-world use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。