arXiv:2602.01047cs.CVcs.AI2026-02中稿 · CVPR被引 5

通过历史信息引导,让大模型更忠于图像内容,减少幻觉。

Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance

  • 利用历史生成记录和内部推理机制修正语言偏见
  • 在多个基准上显著降低物体幻觉,提升视觉定位准确率
  • 无需重新训练,适用于各类大视觉语言模型

大型视觉语言模型(LVLM)能从图文输入中推理并在多种多模态任务中表现良好。尽管如此,它们仍受语言先验影响,常产生幻觉——即语法正确但与视觉输入无关或不匹配的内容。为解决此问题,我们提出残差解码(Residual Decoding, ResDec),一种无需训练的新方法,通过利用历史信息辅助解码。该方法基于LVLM的内在隐式推理机制与词元概率演化机制纠正偏差。大量实验表明,ResDec有效抑制由语言先验引发的幻觉,显著提升视觉定位能力,并减少物体幻觉。此外,该方法在综合性LVLM基准测试中表现优异,彰显其广泛适用性。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) can reason from image-text inputs and perform well in various multimodal tasks. Despite this success, they are affected by language priors and often produce hallucinations. Hallucinations denote generated content that is grammatically and syntactically coherent, yet bears no match or direct relevance to visual input. To address this problem, we propose Residual Decoding (ResDec). It is a novel training-free method that uses historical information to aid decoding. The method relies on the internal implicit reasoning mechanism and token logits evolution mechanism of LVLMs to correct biases. Extensive experiments demonstrate that ResDec effectively suppresses hallucinations induced by language priors, significantly improves visual grounding, and reduces object hallucinations. In addition to mitigating hallucinations, ResDec also performs exceptionally well on comprehensive LVLM benchmarks, highlighting its broad applicability.

视觉语言模型幻觉抑制解码优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。