系统梳理视觉语言模型的攻击策略与防御方法,揭示安全风险
When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs
- 按攻击目标分为越狱、伪装、利用三类,分析数据操纵方法
- 提出攻击分类体系,总结评估指标以量化攻击影响
- 适合关注AI安全、模型鲁棒性的研究者阅读
视觉语言模型(VLMs)因其能有效融合处理文本与视觉信息而备受关注,显著提升了场景理解、机器人等应用性能。然而,其部署也带来重大安全风险,亟需深入研究潜在漏洞。本文全面调研针对VLMs的攻击策略,依据攻击目标(越狱、伪装、利用)进行分类,详述数据操纵方法,并梳理相应防御机制。通过厘清各类攻击间的联系与差异,构建了系统的攻击分类体系,总结涵盖多维度的评估指标。最后,探讨未来研究方向,强调持续探索对提升VLM鲁棒性与安全性的重要性。项目主页持续更新:https://github.com/AobtDai/VLM_Attack_Paper_List。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have gained considerable prominence in recent years due to their remarkable capability to effectively integrate and process both textual and visual information. This integration has significantly enhanced performance across a diverse spectrum of applications, such as scene perception and robotics. However, the deployment of VLMs has also given rise to critical safety and security concerns, necessitating extensive research to assess the potential vulnerabilities these VLM systems may harbor. In this work, we present an in-depth survey of the attack strategies tailored for VLMs. We categorize these attacks based on their underlying objectives - namely jailbreak, camouflage, and exploitation - while also detailing the various methodologies employed for data manipulation of VLMs. Meanwhile, we outline corresponding defense mechanisms that have been proposed to mitigate these vulnerabilities. By discerning key connections and distinctions among the diverse types of attacks, we propose a compelling taxonomy for VLM attacks. Moreover, we summarize the evaluation metrics that comprehensively describe the characteristics and impact of different attacks on VLMs. Finally, we conclude with a discussion of promising future research directions that could further enhance the robustness and safety of VLMs, emphasizing the importance of ongoing exploration in this critical area of study. To facilitate community engagement, we maintain an up-to-date project page, accessible at: https://github.com/AobtDai/VLM_Attack_Paper_List.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。