arXiv:2508.00378cs.AIcs.CV2025-08

通过视觉验证推理过程,减少多模态模型的幻觉问题

CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding

  • 将推理步骤拆解并用图像证据验证每一步
  • 在多个数据集上提升答案准确率和解释可信度
  • 适合关注模型可解释性与可靠性研究者

多模态语言模型在推理时常因仅浅层观察图像而产生幻觉。我们提出CoRGI(Chain-of-Thought Reasoning with Grounded Insights),通过后置视觉验证增强推理可靠性。给定VLM生成的推理过程,CoRGI将其分解为逐步陈述,逐条在图像中寻找视觉依据,并过滤或修正无支撑的表述,再输出最终答案。在五项挑战性基准测试(VCR、ScienceQA、MMMU、MathVista、HallusionBench)上,CoRGI在Qwen-2.5VL、LLaVA-1.6、Gemma3-12B等多个VLM基座上均持续提升答案准确率与解释忠实度。定性分析显示,该验证机制有效降低幻觉并增强可解释性,表明后置视觉接地是构建更可信、透明多模态推理系统的重要方向。

原文摘要 · Abstract (English)

Multimodal reasoning with vision-language models (VLMs) often suffers from hallucinations, as models tend to generate explanations after only a superficial inspection of the image. We present \textbf{CoRGI}(\textbf{C}hain \textbf{o}f \textbf{R}easoning with \textbf{G}rounded \textbf{I}nsights), a framework that enhances reasoning reliability through post-hoc verification of chain-of-thought outputs. Given a VLM-generated rationale, CoRGI decomposes it into step-wise statements, grounds each step in visual evidence, and filters or corrects unsupported claims before producing the final answer. Experiments on five challenging benchmark-VCR, ScienceQA, MMMU, MathVista, and HallusionBenc-demonstrate that CoRGI consistently improves both answer accuracy and explanation faithfulness across multiple VLM backbones, including Qwen-2.5VL, LLaVA-1.6, and Gemma3-12B. Beyond quantitative gains, qualitative analyses further illustrate how the verification process reduces hallucination and strengthens interpretability, suggesting that post-hoc visual grounding is a promising direction for building more trustworthy and transparent multimodal reasoning systems.

多模态推理幻觉抑制可解释性视觉验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。