arXiv:2603.02024cs.CLcs.AI2026-03中稿 · ICLR被引 7

构建真实场景多图推理基准,评估大模型跨图像综合推理能力。

MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning

  • 基于1.9万张真实图像设计多图问答任务,覆盖7类推理类型。
  • 顶尖模型如GPT-5准确率仅58%,且各类推理表现差异大。
  • 揭示思考长度与推理方式对模型表现的影响,适合研究多模态推理者。

多模态大语言模型(MLLMs)在科学分析与数学推理等复杂任务上取得进展,但其在真实生活场景中的跨情境推理能力仍缺乏系统评估。为此,我们提出MMR-Life,一个综合性基准,用于评测MLLMs在真实场景下的多模态多图像推理能力。该基准包含2,646道多选题,基于19,108张主要来自真实世界场景的图像,全面覆盖七类推理:溯因、类比、因果、演绎、归纳、空间和时间推理。与现有基准不同,MMR-Life不依赖领域专长,而是要求模型整合多图信息并应用多样推理技能。对37个先进模型的评估表明,该基准挑战巨大:即使顶级模型如GPT-5准确率也仅为58%,且在不同推理类型间表现波动明显。此外,我们分析了现有MLLMs的推理模式,探讨了思维长度、推理方法与推理类型对性能的影响。总体而言,MMR-Life为下一代多模态推理系统的设计、评估与改进奠定了坚实基础。

原文摘要 · Abstract (English)

Recent progress in the reasoning capabilities of multimodal large language models (MLLMs) has empowered them to address more complex tasks such as scientific analysis and mathematical reasoning. Despite their promise, MLLMs' reasoning abilities across different scenarios in real life remain largely unexplored and lack standardized benchmarks for evaluation. To address this gap, we introduce MMR-Life, a comprehensive benchmark designed to evaluate the diverse multimodal multi-image reasoning capabilities of MLLMs across real-life scenarios. MMR-Life consists of 2,646 multiple-choice questions based on 19,108 images primarily sourced from real-world contexts, comprehensively covering seven reasoning types: abductive, analogical, causal, deductive, inductive, spatial, and temporal. Unlike existing reasoning benchmarks, MMR-Life does not rely on domain-specific expertise but instead requires models to integrate information across multiple images and apply diverse reasoning abilities. The evaluation of 37 advanced models highlights the substantial challenge posed by MMR-Life. Even top models like GPT-5 achieve only 58% accuracy and display considerable variance in performance across reasoning types. Moreover, we analyze the reasoning paradigms of existing MLLMs, exploring how factors such as thinking length, reasoning method, and reasoning type affect their performance. In summary, MMR-Life establishes a comprehensive foundation for evaluating, analyzing, and improving the next generation of multimodal reasoning systems.

多模态推理真实场景多图理解评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。