攻击视觉语言模型的编码视觉令牌,揭示其对视觉特征干扰的脆弱性。
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
- 通过多角度扰动图像编码器输出的视觉令牌,实现非目标攻击。
- 攻击在多个模型和任务间具有强迁移性和泛化能力,成功率超基线方法。
- 为提升视觉语言模型的鲁棒性提供关键测试与改进方向,适合安全研究者参考。
大型视觉语言模型(LVLMs)将视觉信息整合至大语言模型中,展现出出色的多模态对话能力。然而,视觉模块引入了新的鲁棒性挑战:攻击者可生成视觉上无异常的对抗图像,导致模型产生错误回答。现有方法依赖视觉编码器将图像转换为视觉令牌,这些令牌对语言模型理解图像内容至关重要。本文提出一种非目标攻击方法——VT-Attack(视觉令牌攻击),从多维度破坏图像编码器输出的视觉令牌的特征表示、内在关系及语义属性。仅需访问图像编码器,生成的对抗样本在使用相同编码器的不同LVLM之间具有显著迁移性,并在多种任务上表现出强泛化能力。大量实验验证了VT-Attack在攻击性能上优于基线方法,有效揭示了基于视觉编码器的LVLM在视觉特征空间稳定性方面的脆弱性,为模型鲁棒性评估提供了重要依据。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) integrate visual information into large language models, showcasing remarkable multi-modal conversational capabilities. However, the visual modules introduces new challenges in terms of robustness for LVLMs, as attackers can craft adversarial images that are visually clean but may mislead the model to generate incorrect answers. In general, LVLMs rely on vision encoders to transform images into visual tokens, which are crucial for the language models to perceive image contents effectively. Therefore, we are curious about one question: Can LVLMs still generate correct responses when the encoded visual tokens are attacked and disrupting the visual information? To this end, we propose a non-targeted attack method referred to as VT-Attack (Visual Tokens Attack), which constructs adversarial examples from multiple perspectives, with the goal of comprehensively disrupting feature representations and inherent relationships as well as the semantic properties of visual tokens output by image encoders. Using only access to the image encoder in the proposed attack, the generated adversarial examples exhibit transferability across diverse LVLMs utilizing the same image encoder and generality across different tasks. Extensive experiments validate the superior attack performance of the VT-Attack over baseline methods, demonstrating its effectiveness in attacking LVLMs with image encoders, which in turn can provide guidance on the robustness of LVLMs, particularly in terms of the stability of the visual feature space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。