arXiv:2602.15556cs.CV2026-02ACL被引 7

通过挖掘模型内部注意力动态,精准定位关键视觉区域,有效减少多模态模型幻觉。

Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs

  • 利用内部正向注意力动态识别图像核心区域
  • 在多个基准上降低幻觉率,提升视觉定位准确性
  • 无需训练,适合希望改进现有模型的开发者

LVLMs 虽具备强大的多模态推理能力,但仍易产生与视觉输入或用户指令不一致的幻觉。现有无训练方法如对比解码和辅助专家模型常带来数倍计算开销,且可能引入干扰;静态内部信号增强则易受注意力淹没现象影响。我们发现,LVLM 内部的正向注意力动态(PAD)在注意力淹没干扰下仍能自然揭示语义核心视觉区域。基于此,提出无训练注意力干预方法 PADE:构建 PAD 图以识别核心视觉区域,采用每头中位绝对偏差缩放自适应控制干预强度,并通过系统标记补偿维持对复杂指令的关注及长期输出一致性。在多个 LVLM 和基准上的实验表明,PADE 提升了视觉定位能力,减少了幻觉,验证了利用内部注意力动态实现可靠多模态推理的有效性。

原文摘要 · Abstract (English)

LVLMs have achieved strong multimodal reasoning capabilities but remain prone to hallucinations, producing outputs inconsistent with visual inputs or user instructions. Existing training-free methods, including contrastive decoding and auxiliary expert models, which incur several times more computational overhead and may introduce potential interference, as well as static internal signal enhancement, are often vulnerable to the attention sink phenomenon. We find that internal Positive Attention Dynamics (PAD) in LVLMs naturally reveal semantically core visual regions under the distortions of attention sinks. Based on this, we propose Positive Attention Dynamics Enhancement (PADE), a training-free attention intervention that constructs a PAD map to identify semantically core visual regions, applies per-head Median Absolute Deviation Scaling to adaptively control the intervention strength, and leverages System-Token Compensation to maintain attention to complex user instructions and support long-term output consistency. Experiments on multiple LVLMs and benchmarks show that PADE improves visual grounding and reduces hallucinations, validating the effectiveness of leveraging internal attention dynamics for reliable multimodal reasoning.

多模态模型幻觉抑制注意力机制无训练方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。