arXiv:2501.05444cs.CV2025-01ICML被引 152

评测大模型在多模态下的综合推理能力,发现现有模型仍严重不足。

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

  • 构建跨数学、物理、化学、编程的多模态推理任务,要求图文深度融合
  • 顶尖模型在复杂多步任务上表现不佳,链式思维和算力提升效果有限
  • 适合研究多模态推理、模型评估与认知能力提升的研究者参考

人类智能的核心在于能自然地对文字和图像进行综合推理,但当前多模态大语言模型(MLLMs)在该能力上的探索仍不充分。现有基准多偏重文本推理或依赖浅层视觉线索,难以有效评估图文融合的深层推理。本文提出EMMA(Enhanced MultiModal reAsoning),一个面向数学、物理、化学和编程领域的多模态推理基准,任务要求模型必须进行跨模态协同推理,无法通过单一模态独立解决。对主流MLLMs在EMMA上的评估显示,其在复杂多步推理任务中存在显著缺陷,即使采用链式思维提示(Chain-of-Thought)和测试时计算扩展也未能有效提升性能。这表明需发展更优的多模态架构与训练范式,以缩小模型与人类在多模态推理上的差距。

原文摘要 · Abstract (English)

The ability to organically reason over and with both text and images is a pillar of human intelligence, yet the ability of Multimodal Large Language Models (MLLMs) to perform such multimodal reasoning remains under-explored. Existing benchmarks often emphasize text-dominant reasoning or rely on shallow visual cues, failing to adequately assess integrated visual and textual reasoning. We introduce EMMA (Enhanced MultiModal reAsoning), a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding. EMMA tasks demand advanced cross-modal reasoning that cannot be addressed by reasoning independently in each modality, offering an enhanced test suite for MLLMs' reasoning capabilities. Our evaluation of state-of-the-art MLLMs on EMMA reveals significant limitations in handling complex multimodal and multi-step reasoning tasks, even with advanced techniques like Chain-of-Thought prompting and test-time compute scaling underperforming. These findings underscore the need for improved multimodal architectures and training paradigms to close the gap between human and model reasoning in multimodality.

多模态推理模型评估MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。