arXiv:2505.16459cs.AI2025-05被引 3

评测多模态大模型的推理能力,揭示其思维路径缺陷。

MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning

  • 设计包含1083题的多模态推理数据集,覆盖六类高难度推理任务。
  • 顶尖模型如Claude-3.7-Sonnet仍存在思维不一致、过度思考等问题。
  • 提供可量化思维质量的评估流程,适合模型开发者与评测研究者。

多模态大语言模型(MLLMs)已能统一处理语言、视觉和结构化输入,支持逻辑推理、空间分析等复杂任务。然而,尤其是带有中间思维路径的MLLMs-T,其推理能力尚缺乏系统评估。现有工作多关注感知或最终答案正确性,难以洞察模型跨模态的推理过程与失败原因。为此,我们提出MMLU-Reason基准,包含:1)1,083道涵盖六类推理类型的高难度数据集,具有符号深度与多跳需求;2)模块化推理路径评估流水线(RTEP),通过相关性、一致性及结构化错误标注等指标,超越准确率评估推理质量。实验表明,带思维路径的模型整体优于无思维模型,但顶级模型如Claude-3.7-Sonnet与Gemini-2.5 Pro仍存在思维路径不一致、过度推理等病理现象。该基准揭示了准确率与推理质量之间的持续差距,并为未来模型开发提供了可操作的评估框架。MMLU-Reason为下一代多模态推理系统提供了可扩展的评估基础。

原文摘要 · Abstract (English)

Recent advances in Multi-Modal Large Language Models (MLLMs) have enabled unified processing of language, vision, and structured inputs, opening the door to complex tasks such as logical deduction, spatial reasoning, and scientific analysis. Despite their promise, the reasoning capabilities of MLLMs, particularly those augmented with intermediate thinking traces (MLLMs-T), remain poorly understood and lack standardized evaluation benchmarks. Existing work focuses primarily on perception or final answer correctness, offering limited insight into how models reason or fail across modalities. To address this gap, we introduce the MMLU-Reason, a new benchmark designed to rigorously evaluate multi-modal reasoning with explicit thinking. The MMLU-Reason comprises 1) a high-difficulty dataset of 1,083 questions spanning six diverse reasoning types with symbolic depth and multi-hop demands and 2) a modular Reasoning Trace Evaluation Pipeline (RTEP) for assessing reasoning quality beyond accuracy through metrics like relevance, consistency, and structured error annotations. Empirical results show that MLLMs-T overall outperform non-thinking counterparts, but even top models like Claude-3.7-Sonnet and Gemini-2.5 Pro suffer from reasoning pathologies such as inconsistency and overthinking. This benchmark reveals persistent gaps between accuracy and reasoning quality and provides an actionable evaluation pipeline for future model development. Overall, the MMLU-Reason offers a scalable foundation for evaluating, comparing, and improving the next generation of multi-modal reasoning systems.

多模态推理评测大模型思维路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。