arXiv:2505.24238cs.CVcs.LG2025-05NeurIPS被引 23

提出新基准与方法,精准识别并减少多模态模型的推理幻觉。

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

  • 构建隔离推理幻觉的测试集,区分感知与推理错误。
  • 发现模型规模和训练阶段显著影响幻觉程度,空间关系误判仍严重。
  • 提出渐进式强化微调与协作提示推理,有效降低逻辑幻觉。

多模态大语言模型(MLLM)中的多源幻觉限制了其推理正确性。现有基准难以区分感知引发的幻觉与推理引发的幻觉,阻碍了对模型推理失败的诊断。为此,我们提出 { extit{dataset}} 基准,通过构造图像被正确感知但推理仍出错的问题,专门评估推理幻觉。该基准引入多粒度评估指标:准确率、事实性及大模型幻觉得分,实现幻觉量化。分析发现:(1)模型规模、数据规模和训练阶段显著影响逻辑、虚构和事实性幻觉;(2)当前MLLM在空间关系误判导致的空间幻觉上无明显改善,显示其视觉推理能力有限;(3)不同题型对应不同幻觉模式,揭示特定挑战与缓解方向。为此,我们提出 { extit{method}},结合课程式强化微调(逐步降低学习难度以生成逻辑一致的推理链)与协作提示推理(降低推理复杂度)。{ extit{method}} 在 { extit{dataset}} 上建立基线,有效减少原始基础模型的逻辑幻觉。

原文摘要 · Abstract (English)

Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to adequately distinguish between perception-induced hallucinations and reasoning-induced hallucinations. This failure constitutes a significant issue and hinders the diagnosis of multimodal reasoning failures within MLLMs. To address this, we propose the {\dataset} benchmark, which isolates reasoning hallucinations by constructing questions where input images are correctly perceived by MLLMs yet reasoning errors persist. {\dataset} introduces multi-granular evaluation metrics: accuracy, factuality, and LLMs hallucination score for hallucination quantification. Our analysis reveals that (1) the model scale, data scale, and training stages significantly affect the degree of logical, fabrication, and factual hallucinations; (2) current MLLMs show no effective improvement on spatial hallucinations caused by misinterpreted spatial relationships, indicating their limited visual reasoning capabilities; and (3) question types correlate with distinct hallucination patterns, highlighting targeted challenges and potential mitigation strategies. To address these challenges, we propose {\method}, a method that combines curriculum reinforcement fine-tuning to encourage models to generate logic-consistent reasoning chains by stepwise reducing learning difficulty, and collaborative hint inference to reduce reasoning complexity. {\method} establishes a baseline on {\dataset}, and reduces the logical hallucinations in original base models.

多模态幻觉检测推理链评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。