构建首个覆盖三类逻辑推理的多模态大模型评测基准
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
- 设计涵盖归纳、演绎、溯因三类推理的题目体系
- 顶尖模型在综合推理测试中表现有限且类型间差异显著
- 揭示思维模式与规则强化对推理提升的局限性
逻辑推理是人类智能的核心能力,也是多模态大语言模型(MLLMs)的关键需求。尽管多模态推理取得进展,现有评测基准因缺乏对逻辑推理类型的明确定义和理解,难以全面评估模型能力。为此,我们提出MME-Reasoning,一个全面评估MLLM推理能力的基准,覆盖归纳、演绎和溯因三类推理。通过精心筛选数据,确保题目聚焦推理能力而非感知或知识广度,并扩展评估协议以涵盖多样化问题。评估显示,当前最先进MLLM在整体逻辑推理能力上存在显著不足,且不同推理类型间表现不平衡。此外,我们深入分析了常被认为能提升推理能力的‘思维模式’和基于规则的强化学习方法,结果表明其效果有限。这些发现揭示了当前MLLM在多种逻辑推理场景下的关键短板,为系统理解与评估推理能力提供了全面洞察。
原文摘要 · Abstract (English)
Logical reasoning is a fundamental aspect of human intelligence and an essential capability for multimodal large language models (MLLMs). Despite the significant advancement in multimodal reasoning, existing benchmarks fail to comprehensively evaluate their reasoning abilities due to the lack of explicit categorization for logical reasoning types and an unclear understanding of reasoning. To address these issues, we introduce MME-Reasoning, a comprehensive benchmark designed to evaluate the reasoning ability of MLLMs, which covers all three types of reasoning (i.e., inductive, deductive, and abductive) in its questions. We carefully curate the data to ensure that each question effectively evaluates reasoning ability rather than perceptual skills or knowledge breadth, and extend the evaluation protocols to cover the evaluation of diverse questions. Our evaluation reveals substantial limitations of state-of-the-art MLLMs when subjected to holistic assessments of logical reasoning capabilities. Even the most advanced MLLMs show limited performance in comprehensive logical reasoning, with notable performance imbalances across reasoning types. In addition, we conducted an in-depth analysis of approaches such as ``thinking mode'' and Rule-based RL, which are commonly believed to enhance reasoning abilities. These findings highlight the critical limitations and performance imbalances of current MLLMs in diverse logical reasoning scenarios, providing comprehensive and systematic insights into the understanding and evaluation of reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。