arXiv:2512.05546cs.CVcs.AI2025-12

通过感知视觉-文本协同度,动态修正模型注意力,减少幻觉。

Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models

  • 用博弈论方法实时检测视觉与语言的协同程度
  • 在推理时选择性重定向中间层注意力,防止脱离视觉证据
  • 无需训练即可通用适配多模型,适合提升生成可靠性

大视觉语言模型常出现文本惯性,即注意力从视觉证据偏移至语言先验,导致物体幻觉。现有解码策略仅在输出层干预,无法纠正内部推理偏差;近期基于启发式头抑制或全局控制向量的方法缺乏理论依据。本文提出无需训练、推理时使用的 Conscious Gaze (CG-VLM) 框架,将博弈论可解释性转化为可操作的解码控制。基于 Harsanyi 交互构建的认知需求传感器,实时估计视觉-文本协同度,识别需强化视觉引导的时刻。在此信号下,聚焦共识诱导模块在注意力坍缩前,选择性重定向中间层注意力至视觉标记。CG-VLM 在 POPE 与 CHAIR 多个模型(InstructBLIP、LLaVA、Qwen-VL、mPLUG)上均达当前最优表现,同时保持通用能力,证明了基于令牌级感知的精准、上下文敏感干预可行,且不损害基础知识。

原文摘要 · Abstract (English)

Large Vision-Language Models (VLMs) often exhibit text inertia, where attention drifts from visual evidence toward linguistic priors, resulting in object hallucinations. Existing decoding strategies intervene only at the output logits and thus cannot correct internal reasoning drift, while recent internal-control methods based on heuristic head suppression or global steering vectors lack principled grounding. We introduce Conscious Gaze (CG-VLM), a training-free, inference-time framework that converts game-theoretic interpretability into actionable decoding control. A Cognitive Demand Sensor built on Harsanyi interactions estimates instantaneous vision-text synergy and identifies moments when visual grounding is necessary. Conditioned on this signal, a Focused Consensus Induction module selectively reorients mid-layer attention toward visual tokens before collapse into text priors. CG-VLM achieves state-of-the-art results on POPE and CHAIR across InstructBLIP, LLaVA, Qwen-VL, and mPLUG, while preserving general capabilities, demonstrating that token-level sensing enables precise, context-aware intervention without compromising foundational knowledge.

视觉语言模型注意力机制幻觉抑制推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。