通过视觉注意力引导双锚点解码,减少多模态模型幻觉
Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation
- 用视觉注意力选关键层,动态校准每步生成
- 在多个基准上显著降低幻觉率,提升推理能力
- 适合关注多模态生成准确性的研究者
多模态大语言模型虽具备强大推理能力,但仍存在生成内容与视觉内容矛盾的幻觉问题。本文提出双锚点自省解码(DaID),一种对比解码框架,通过挖掘模型内部感知差异来动态校准每个词元生成。具体而言,DaID识别出一个亮点层以增强视觉事实信号,一个暗影层以抑制文本惯性。利用视觉注意力分布引导这一双锚点选择过程,确保精准的、逐词元的适应性调整。在多个基准和多模态大模型上的实验表明,DaID能显著缓解幻觉问题,同时提升通用推理能力。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities yet continue to suffer from hallucination, where generated text contradicts visual content. In this paper, we introduce Dual-Anchor Introspective Decoding (DaID), a novel contrastive decoding framework that dynamically calibrates each token generation by mining the model's internal perceptual discrepancies. Specifically, DaID identifies a Spotlight layer to amplify visual factual signals and a Shadow layer to suppress textual inertia. By leveraging visual attention distributions to guide this dual-anchor selection process, our method ensures precise, token-specific adaptation. Experimental results across multiple benchmarks and MLLMs demonstrate that DaID significantly mitigates hallucination while enhancing general reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。