arXiv:2601.22735cs.CL2026-01被引 1

评测视觉大模型推理时的幻觉问题,揭示思考过程如何影响判断准确性。

MM-THEBench: Do Reasoning MLLMs Think Reasonably?

  • 构建细粒度认知维度的幻觉分类体系,评估推理中间步骤
  • 发现自我反思虽增强鲁棒性,却也引入新幻觉,微小感知错误仍致误判
  • 适合关注多模态模型可靠性与推理机制的研究者使用

近期多模态大语言模型(MLLMs)的发展标志着从非思考模型向后训练推理模型的转变,能够通过思考解决复杂问题。然而,这种思考是否能缓解多模态感知与推理中的幻觉尚不明确。自我反思推理虽提升鲁棒性,却引入额外幻觉,细微感知错误仍导致错误或偶然正确的答案。现有基准主要针对推理前模型,忽视内部思考过程,无法衡量思考中产生的幻觉。为此,我们提出MM-THEBench,一个全面评估推理型MLLM中间思维链(CoTs)幻觉的基准。该基准包含基于认知维度的细粒度分类体系、经验证的推理标注数据及多层级自动化评估框架。对主流推理型MLLM的广泛实验揭示了思考对幻觉与推理能力在多种多模态任务中的影响。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) mark a shift from non-thinking models to post-trained reasoning models capable of solving complex problems through thinking. However, whether such thinking mitigates hallucinations in multimodal perception and reasoning remains unclear. Self-reflective reasoning enhances robustness but introduces additional hallucinations, and subtle perceptual errors still result in incorrect or coincidentally correct answers. Existing benchmarks primarily focus on models before the emergence of reasoning MLLMs, neglecting the internal thinking process and failing to measure the hallucinations that occur during thinking. To address these challenges, we introduce MM-THEBench, a comprehensive benchmark for assessing hallucinations of intermediate CoTs in reasoning MLLMs. MM-THEBench features a fine-grained taxonomy grounded in cognitive dimensions, diverse data with verified reasoning annotations, and a multi-level automated evaluation framework. Extensive experiments on mainstream reasoning MLLMs reveal insights into how thinking affects hallucination and reasoning capability in various multimodal tasks.

多模态推理模型幻觉检测评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。