让视觉语言模型在不同问题下都输出指定答案,提升攻击鲁棒性。
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models

- 通过操控注意力分布,让图像信息主导生成结果。
- 在多种查询下仍能保持攻击效果,跨查询成功率显著提升。
- 适合研究模型安全与注意力机制的学者参考。
现有针对视觉语言模型(VLMs)的对抗攻击在面对不同文本查询时,效果常会下降。本文研究跨查询响应操纵问题,即同一扰动输入需对多种用户查询保持有效。分析发现,成功迁移与生成过程中维持图像主导的注意力模式密切相关。为此,提出新攻击方法「Attention Hijacking」,显式引导内部注意力分布趋向持久的图像主导模式,增强视觉标记对目标输出标记的影响,同时抑制文本标记的竞争影响,从而降低输出对查询措辞的依赖。大量实验表明,该方法在多个主流VLM上显著提升了跨查询可迁移性,适用于多种攻击场景,为理解注意力稳定性在可迁移响应操纵中的作用提供了新见解。
原文摘要 · Abstract (English)
Existing adversarial attacks on vision-language models (VLMs) can steer model outputs toward attacker-specified target responses, but their effectiveness often degrades when the same perturbed input is paired with different textual queries. This paper studies cross-query response manipulation, where a single adversarial example is expected to remain effective across diverse user queries. We first analyze the limitations of existing attacks and find that successful transfer is closely associated with preserving an image-dominant attention pattern during response generation. Motivated by the observation, we propose \textbf{Attention Hijacking}, a novel adversarial attack that explicitly steers internal attention distributions toward a persistent image-dominant pattern. By amplifying the influence of visual tokens on target response tokens while suppressing the competing influence of textual tokens, our method reduces the dependence of the manipulated output on the specific wording of the query. Extensive experiments on widely used VLMs show that Attention Hijacking substantially improves cross-query transferability across diverse target responses and unseen queries. The method also extends effectively to multiple attack scenarios, offering new insights into the role of attention stability in transferable response manipulation for VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。