arXiv:2510.06292cs.CVcs.AI2025-10中稿 · ICLR被引 1

通过图文交错推理链,减少视觉语言模型的关系幻觉。

ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations

  • 用多视角问题引导图文交替推理,逐步深化关系理解。
  • 在多个基准上显著降低关系幻觉率,提升推理可靠性。
  • 无需训练,适合希望提升现有模型可信度的研究者。

尽管大型视觉语言模型(LVLMs)在多模态任务中表现优异,但幻觉问题仍严重影响其可靠性。其中,关系幻觉占比最高却最被忽视。为此,本文提出ChainMPQ(多视角问题引导的图文交错推理链),一种无需训练的方法,通过累积的文本与视觉记忆改进LVLM的关系推理能力。首先从问题中提取主语和宾语关键词,增强图像对应区域;随后构建聚焦于关系三要素(主语、宾语、关联)的多视角问题,按序输入模型,前序步骤的图文记忆为后续提供上下文支持,形成交错推理链,实现渐进式关系推断。在多个LVLM和基准上的实验表明,ChainMPQ显著降低关系幻觉,消融实验证明其三个核心模块均有效。

原文摘要 · Abstract (English)

While Large Vision-Language Models (LVLMs) achieve strong performance in multimodal tasks, hallucinations continue to hinder their reliability. Among the three categories of hallucinations, which include object, attribute, and relation, relation hallucinations account for the largest proportion but have received the least attention. To address this issue, we propose ChainMPQ (Multi-Perspective Questions guided Interleaved Chain of Image and Text), a training-free method that improves relational inference in LVLMs by utilizing accumulated textual and visual memories. ChainMPQ first extracts subject and object keywords from the question to enhance the corresponding image regions. It then constructs multi-perspective questions that focus on the three core components of a relationship: the subject, the object, and the relation that links them. These questions are sequentially input to the model, with textual and visual memories from earlier steps providing supporting context for subsequent ones, thereby forming an interleaved chain of images and text that guides progressive relational reasoning. Experiments on multiple LVLMs and benchmarks show that ChainMPQ substantially reduces relation hallucinations, while ablation studies further validate the effectiveness of its three core modules.

视觉语言模型关系推理幻觉抑制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。