arXiv:2604.01989cs.CVcs.AI2026-04

破解视觉注意力僵化,提升多模态模型推理能力

Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation

  • 通过动态激发新出现的视觉特征,打破注意力固化模式
  • 在多个基准上显著降低认知幻觉,无需额外训练
  • 适合需要精准关系推理的多模态应用开发者

与物体静止则保持静止类似,我们发现多模态大语言模型(MLLMs)中的视觉注意力在早期解码阶段趋于稳定后会表现出明显惯性,难以动态支持组合式理解所需的认知推理。现有幻觉缓解方法主要针对对象存在性或属性等感知幻觉,对需跨对象关系推断的认知幻觉仍不充分。通过逐令牌分析,我们识别出视觉惯性是关键因素:语义关键区域的注意力持续集中,无法动态响应关系推理需求。为此,我们提出惯性感知视觉激励(IVE),将认知推理建模为视觉注意力的动态响应。具体地,IVE选择相对于历史注意力趋势动态浮现的视觉令牌,同时区分惯性行为。为进一步促进组合推理,IVE引入惯性感知惩罚,抑制过度聚焦并限制注意力在局部区域的持久性。大量实验表明,IVE在不增加训练成本的前提下,有效提升多种MLLM在多个基准上的表现。

原文摘要 · Abstract (English)

Like a body at rest that stays at rest, we find that visual attention in multimodal large language models (MLLMs) exhibits pronounced inertia, remaining largely static once settled during early decoding steps and failing to support the compositional understanding required for cognitive inference. While existing hallucination mitigation methods mainly target perceptual hallucinations concerning object existence or attributes, they remain inadequate for such cognitive hallucinations that require inter-object relational deduction. Through token-wise analysis, we identify visual inertia as a contributing factor: attention to semantically critical regions remains persistently focused and fails to dynamically support relational inference. We thereby propose Inertia-aware Visual Excitation (IVE) that breaks this inertial pattern by modeling cognitive inference as the dynamic responsiveness of visual attention. Specifically, IVE selects visual tokens that are dynamically emerging relative to historical attention trends while distinguishing tokens exhibiting inertial behavior. To further facilitate compositional inference, IVE introduces an inertia-aware penalty that discourages over-concentration and limits the persistence of attention within localized regions. Extensive experiments show the effectiveness of IVE across various MLLMs and benchmarks without additional training.

多模态注意力机制幻觉缓解推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。