arXiv:2505.19015cs.CVcs.MM2025-05ACL被引 30

新基准测试揭示大模型空间关系理解能力严重不足

Can Multimodal Large Language Models Understand Spatial Relations?

  • 构建人工标注的时空关系数据集SpatialMQA,避免依赖边界框和先验知识
  • 当前顶尖多模态模型仅48.14%准确率,远低于人类98.40%水平
  • 适合研究视觉-语言对齐、认知推理与评估方法的学者使用

空间关系推理是多模态大模型理解客观世界的关键任务。现有基准存在依赖边界框、忽略视角变化或仅靠模型先验知识即可作答等问题。为此,我们提出了基于COCO2017的人工标注空间关系推理基准SpatialMQA,使模型更关注图像本身理解。通过定制化标注流程,获得5,392个高质量样本。在此基础上,对一系列闭源与开源多模态大模型进行评估,结果显示当前最先进模型准确率仅为48.14%,远低于人类水平(98.40%)。大量实验分析揭示了未来研究方向。数据集与代码已公开于https://github.com/ziyan-xiaoyu/SpatialMQA.git。

原文摘要 · Abstract (English)

Spatial relation reasoning is a crucial task for multimodal large language models (MLLMs) to understand the objective world. However, current benchmarks have issues like relying on bounding boxes, ignoring perspective substitutions, or allowing questions to be answered using only the model's prior knowledge without image understanding. To address these issues, we introduce SpatialMQA, a human-annotated spatial relation reasoning benchmark based on COCO2017, which enables MLLMs to focus more on understanding images in the objective world. To ensure data quality, we design a well-tailored annotation procedure, resulting in SpatialMQA consisting of 5,392 samples. Based on this benchmark, a series of closed- and open-source MLLMs are implemented and the results indicate that the current state-of-the-art MLLM achieves only 48.14% accuracy, far below the human-level accuracy of 98.40%. Extensive experimental analyses are also conducted, suggesting the future research directions. The benchmark and codes are available at https://github.com/ziyan-xiaoyu/SpatialMQA.git.

多模态空间推理评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。