arXiv:2603.25088cs.CV2026-03

通过中间层视觉锚点,减少多模态大模型的幻觉问题

Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors

  • 用中间层视觉特征作为锚点,抑制深层注意力回退到早期噪声
  • 在多个模型和数据集上显著降低幻觉率,计算开销几乎不变
  • 无需训练,适合快速部署到现有多模态大模型中

多模态大语言模型常出现物体幻觉。现有研究虽采用注意力增强和视觉重追踪,但对模型深层注意力漂移的解释不足。本文分析了视觉特征在各层的演化过程,发现幻觉源于深层注意力逐渐回归到早期层的初始噪声。我们观察到输出可靠性依赖于中间层获取的视觉锚点,而非最终层。基于此,提出CLVA(跨层视觉锚点)方法,无需训练,强化关键中间层特征并抑制回退噪声。该方法通过利用注意力动态捕获的锚点,有效将深层注意力拉回正确视觉区域。我们在多种架构与基准上验证,性能优异,且计算时间与显存占用无显著增加。

原文摘要 · Abstract (English)

Multimodal Large Language Models often suffer from object hallucination. While existing research utilizes attention enhancement and visual retracing, we find these works lack sufficient interpretability regarding attention drift in final model stages. In this paper, we investigate the layer wise evolution of visual features and discover that hallucination stems from deep layer attention regressing toward initial visual noise from early layers. We observe that output reliability depends on acquiring visual anchors at intermediate layers rather than final layers. Based on these insights, we propose CLVA, which stands for Cross-Layer Visual Anchors, a training free method that reinforces critical mid layer features while suppressing regressive noise. This approach effectively pulls deep layer attention back to correct visual regions by utilizing essential anchors captured from attention dynamics. We evaluate our method across diverse architectures and benchmarks, demonstrating outstanding performance without significant increase in computational time and GPU memory.

多模态幻觉抑制视觉锚点注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。