arXiv:2506.23563cs.AIcs.CL2025-06ICCV被引 18

构建开放式多模态长链推理评测集,精准评估大模型逻辑能力。

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

  • 从6大学科、多难度层级设计需多步推理的开放问题
  • 用多模型投票过滤猜测与记忆类题目,提升评测可靠性
  • 标注步骤解并采用三元评分法,可评估中间推理过程

推理在推动多模态大语言模型(MLLM)向通用人工智能发展方面至关重要。然而,现有MLLM评测基准在三个关键方面存在不足:(1) 难度与多样性不足,(2) 易受猜测和记忆干扰,(3) 对中间推理步骤评估不充分。为此,我们提出MMReason,一个专为精准、全面评估MLLM长链推理能力而设计的新基准。首先,从6个学科领域、多个难度级别(从中学到大学,从基础到竞赛级)收集需多步推理的挑战性问题。其次,将问题转化为开放格式,并通过多模型投票机制筛选,剔除依赖猜测或记忆的捷径案例,确保评测稳健性。第三,对问题标注详细分步解答,并设计基于参考答案的三元评分机制,可靠评估中间推理步骤。我们利用MMReason对主流先进MLLM进行评测,并深入分析其推理表现。代码将公开于https://github.com/HJYao00/MMReason。

原文摘要 · Abstract (English)

Reasoning plays a crucial role in advancing Multimodal Large Language Models (MLLMs) toward Artificial General Intelligence. However, existing MLLM benchmarks often fall short in precisely and comprehensively evaluating long-chain reasoning abilities from three key aspects: (1) lack of difficulty and diversity, (2) susceptibility to guessability and memorization, (3) inadequate assessment of intermediate reasoning steps. To fill this gap, we introduce MMReason, a new benchmark designed to precisely and comprehensively evaluate MLLM long-chain reasoning capability with diverse, open-ended, challenging questions. First, we curate challenging questions requiring multi-step reasoning from various fields (i.e., 6 disciplines) and multiple difficulty levels (i.e., from pre-university to university, and from foundational to competition tiers). Second, these questions are reformulated into an open-ended format and filtered using a multi-model voting technique to eliminate shortcut cases related to guessing and memorization, ensuring robust reasoning evaluations. Third, we annotate the questions with detailed step-by-step solutions, and design a reference-based ternary scoring mechanism to reliably assess intermediate reasoning steps. With MMReason, we benchmark popular leading MLLMs and provide an in-depth analysis of their reasoning capabilities. We hope MMReason will serve as a valuable resource for advancing MLLM reasoning research. Code will be available at https://github.com/HJYao00/MMReason.

多模态推理评测基准长链思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。