arXiv:2509.00371cs.CV2025-09被引 1

区分遗漏与虚构幻觉根源,提出新方法减少遗漏不增加虚构。

Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs

  • 区分遗漏与虚构幻觉的成因:信心不足与训练数据偏差。
  • 提出VPFC方法,在不增加虚构的前提下有效减少遗漏幻觉。
  • 适合关注多模态模型可靠性与幻觉机制的研究者。

多模态大语言模型虽取得显著进展,但物体幻觉仍是持续挑战。现有方法基于错误假设,认为遗漏与虚构幻觉源于同一原因,常导致减少遗漏时引发更多虚构。本文通过视觉注意力干预实验发现:遗漏幻觉源于将视觉特征映射为语言表达时信心不足,而虚构幻觉则源于跨模态表示空间中的虚假关联,由训练语料的统计偏差引起。基于此,提出视觉-语义注意力势场框架,揭示模型如何构建视觉证据以推断物体存在与否。据此设计可即插即用的视觉势场校准(VPFC)方法,有效降低遗漏幻觉,且不引入额外虚构幻觉。研究揭示当前幻觉研究的关键疏漏,为发展更鲁棒、平衡的幻觉缓解策略指明新方向。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved impressive advances, yet object hallucination remains a persistent challenge. Existing methods, based on the flawed assumption that omission and fabrication hallucinations share a common cause, often reduce omissions only to trigger more fabrications. In this work, we overturn this view by demonstrating that omission hallucinations arise from insufficient confidence when mapping perceived visual features to linguistic expressions, whereas fabrication hallucinations result from spurious associations within the cross-modal representation space due to statistical biases in the training corpus. Building on findings from visual attention intervention experiments, we propose the Visual-Semantic Attention Potential Field, a conceptual framework that reveals how the model constructs visual evidence to infer the presence or absence of objects. Leveraging this insight, we introduce Visual Potential Field Calibration (VPFC), a plug-and-play hallucination mitigation method that effectively reduces omission hallucinations without introducing additional fabrication hallucinations. Our findings reveal a critical oversight in current object hallucination research and chart new directions for developing more robust and balanced hallucination mitigation strategies.

多模态幻觉机制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。