arXiv:2505.23764cs.CVcs.CL2025-05被引 175

评测多图空间推理能力,揭示大模型在复杂场景理解上的巨大差距

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

  • 基于12万张图像构建千道多图问答题,每题含逐步推理过程
  • 最强开源模型仅30%准确率,人类达97%,凸显模型短板
  • 提供错误分析工具,定位四类典型空间推理失败模式

空间智能对在复杂物理世界中运行的多模态大语言模型至关重要。现有基准仅考察单图关系,无法评估真实场景所需的多图空间推理能力。我们提出MMSI-Bench,一个专注于多图空间智能的视觉问答基准。六位3D视觉研究员耗时超过300小时,从12万余张图像中精心设计了1000道高难度、无歧义的多项选择题,每题配有精心设计的干扰项和分步推理流程。我们对37个开源及专有MLLM进行了广泛实验,结果显示显著差距:最强开源模型准确率约30%,OpenAI GPT-5推理模型达40%,而人类得分高达97%。这些结果凸显了MMSI-Bench的挑战性及未来研究的巨大提升空间。利用标注的推理过程,我们还提供了自动化错误分析流水线,诊断出四大主要失败模式:(1) 空间定位错误,(2) 重叠匹配与场景重建错误,(3) 情境转换推理错误,(4) 空间逻辑错误,为推进空间智能提供了洞见。

原文摘要 · Abstract (English)

Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relations and thus fail to assess the multi-image spatial reasoning that real-world deployments demand. We introduce MMSI-Bench, a VQA benchmark dedicated to multi-image spatial intelligence. Six 3D-vision researchers spent more than 300 hours meticulously crafting 1,000 challenging, unambiguous multiple-choice questions from over 120,000 images, each paired with carefully designed distractors and a stepwise reasoning process. We conduct extensive experiments and evaluate 37 open-source and proprietary MLLMs, observing a wide gap: the strongest open-source model attains roughly 30% accuracy and OpenAI's GPT-5 reasoning model reaches 40%, while humans score 97%. These results underscore the challenging nature of MMSI-Bench and the substantial headroom for future research. Leveraging the annotated reasoning processes, we also provide an automated error analysis pipeline that diagnoses four dominant failure modes, including (1) grounding errors, (2) overlap-matching and scene-reconstruction errors, (3) situation-transformation reasoning errors, and (4) spatial-logic errors, offering insights for advancing spatial intelligence. Project page: https://runsenxu.com/projects/MMSI_Bench .

空间推理多图理解大模型评测视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。