图像微扰可精准操控视觉语言模型输出,存在安全风险与保护潜力。
Attention! Your Vision Language Model Could Be Maliciously Manipulated
- 通过梯度优化与可微变换生成不可察觉的图像扰动
- 攻击可精准控制每个输出词元,实现越狱、劫持等多类恶意行为
- 兼具攻击与版权水印功能,适用于模型安全性评估
大型视觉语言模型(VLMs)在理解复杂现实场景和支撑数据驱动决策方面取得显著进展。然而,这些模型对文本或图像类对抗样本表现出严重脆弱性,可能导致越狱、劫持、幻觉等恶意后果。本文通过实证与理论分析表明,VLMs尤其易受基于图像的对抗样本攻击,微小扰动即可精确操控每个输出词元。为此,我们提出新型攻击方法——视觉-语言模型操纵攻击(VMA),融合一阶与二阶动量优化技术,并结合可微变换机制,有效优化对抗扰动。值得注意的是,VMA具有双重用途:既可用于实施越狱、劫持、隐私泄露、拒绝服务及生成海绵样样本等攻击,也可用于注入版权水印以实现保护。大量实证评估验证了VMA在多种场景与数据集上的有效性与泛化能力。代码已开源:https://github.com/Trustworthy-AI-Group/VMA。
原文摘要 · Abstract (English)
Large Vision-Language Models (VLMs) have achieved remarkable success in understanding complex real-world scenarios and supporting data-driven decision-making processes. However, VLMs exhibit significant vulnerability against adversarial examples, either text or image, which can lead to various adversarial outcomes, e.g., jailbreaking, hijacking, and hallucination, etc. In this work, we empirically and theoretically demonstrate that VLMs are particularly susceptible to image-based adversarial examples, where imperceptible perturbations can precisely manipulate each output token. To this end, we propose a novel attack called Vision-language model Manipulation Attack (VMA), which integrates first-order and second-order momentum optimization techniques with a differentiable transformation mechanism to effectively optimize the adversarial perturbation. Notably, VMA can be a double-edged sword: it can be leveraged to implement various attacks, such as jailbreaking, hijacking, privacy breaches, Denial-of-Service, and the generation of sponge examples, etc, while simultaneously enabling the injection of watermarks for copyright protection. Extensive empirical evaluations substantiate the efficacy and generalizability of VMA across diverse scenarios and datasets. Code is available at https://github.com/Trustworthy-AI-Group/VMA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。