arXiv:2503.08216cs.CV2025-03被引 13

发现视觉语言模型幻觉源于指令令牌干扰,提出无训练方法解耦干扰。

Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation

  • 通过计算指令驱动的视觉显著性,识别干扰视觉注意力的指令令牌。
  • 掩蔽干扰令牌的视觉注意力,显著降低多个基准上的幻觉率。
  • 适合关注模型可解释性与幻觉缓解的研究者和开发者。

尽管大型视觉语言模型(LVLMs)取得成功,但仍易产生幻觉。现有研究多归因于对图像标记的视觉注意力不足,但我们的发现表明,幻觉还源于解码过程中指令标记的干扰。某些指令标记会持续扭曲模型的视觉感知,劫持其对不具判别性的视觉区域的注意力,阻碍模型整合图像的全局上下文信息,最终导致幻觉。我们称此现象为‘注意力劫持’,相应干扰标记为‘注意力劫持者’。为此,我们提出无需训练的新策略——注意力劫持者检测与解耦(AID),旨在分离劫持者的影响,使模型能依赖其内在的上下文感知注意力图。AID包含三部分:首先,通过计算指令驱动的视觉显著性识别注意力劫持者;其次,设计注意力解耦机制掩蔽这些劫持者的视觉注意力,减轻其对后续标记的干扰;最后,重新平衡指令驱动与图像驱动的视觉显著性,避免过度掩蔽。大量实验表明,AID在多个基准上显著减少各类LVLM的幻觉。

原文摘要 · Abstract (English)

Despite their success, Large Vision-Language Models (LVLMs) remain vulnerable to hallucinations. While existing studies attribute the cause of hallucinations to insufficient visual attention to image tokens, our findings indicate that hallucinations also arise from interference from instruction tokens during decoding. Intuitively, certain instruction tokens continuously distort LVLMs' visual perception during decoding, hijacking their visual attention toward less discriminative visual regions. This distortion prevents them integrating broader contextual information from images, ultimately leading to hallucinations. We term this phenomenon 'Attention Hijacking', where disruptive instruction tokens act as 'Attention Hijackers'. To address this, we propose a novel, training-free strategy namely Attention HIjackers Detection and Disentanglement (AID), designed to isolate the influence of Hijackers, enabling LVLMs to rely on their context-aware intrinsic attention map. Specifically, AID consists of three components: First, Attention Hijackers Detection identifies Attention Hijackers by calculating instruction-driven visual salience. Next, Attention Disentanglement mechanism is proposed to mask the visual attention of these identified Hijackers, and thereby mitigate their disruptive influence on subsequent tokens. Finally, Re-Disentanglement recalculates the balance between instruction-driven and image-driven visual salience to avoid over-masking effects. Extensive experiments demonstrate that AID significantly reduces hallucination across various LVLMs on several benchmarks.

视觉语言模型幻觉缓解注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。