arXiv:2506.04280cs.CVcs.AI2025-06被引 19

首个评估多图推理能力的基准,揭示开源模型差距。

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

  • 构建首个支持结构化多图推理的评测基准
  • 40个模型测试显示开源模型显著落后于商业模型
  • 适合研究多模态推理与奖励模型的学者使用

随着多模态大语言模型(MLLMs)能力提升和广泛应用,其同时处理和推理多张图像的需求日益增长。然而现有基准大多聚焦单图视觉推理或仅以最终答案评估多图理解任务,忽视了对多图输入下推理能力的系统考察。为此,我们提出首个面向多图结构化视觉推理的基准——多模态多图推理基准(MMRB),包含92个子任务,覆盖空间、时间与语义推理,所有任务均配有由GPT-4o生成并经人工精修的多解、带思维链(CoT)标注。衍生子集用于评估多模态奖励模型在多图场景下的表现。为实现快速可扩展评估,我们设计基于开源LLM的句子级匹配框架。对40个MLLM的广泛实验表明,开源模型在多图推理任务中仍显著落后于商业模型;当前多模态奖励模型几乎无法完成多图排序任务。

原文摘要 · Abstract (English)

With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously. However, existing MLLM benchmarks focus either on single-image visual reasoning or on multi-image understanding tasks with only final-answer evaluation, leaving the reasoning capabilities of MLLMs over multi-image inputs largely underexplored. To address this gap, we introduce the $\textbf{Multimodal Multi-image Reasoning Benchmark (MMRB)}$, the first benchmark designed to evaluate structured visual reasoning across multiple images. MMRB comprises $\textbf{92 sub-tasks}$ covering spatial, temporal, and semantic reasoning, with multi-solution, CoT-style annotations generated by GPT-4o and refined by human experts. A derivative subset is designed to evaluate multimodal reward models in multi-image scenarios. To support fast and scalable evaluation, we propose a sentence-level matching framework using open-source LLMs. Extensive baseline experiments on $\textbf{40 MLLMs}$, including 9 reasoning-specific models and 8 reward models, demonstrate that open-source MLLMs still lag significantly behind commercial MLLMs in multi-image reasoning tasks. Furthermore, current multimodal reward models are nearly incapable of handling multi-image reward ranking tasks.

多模态推理评测图像理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。