arXiv:2505.17012cs.CVcs.AI2025-05中稿 · CVPR被引 9

构建首个覆盖30类任务的多模态空间智能评测基准

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

  • 设计涵盖多种视觉数据与问答形式的综合评测集
  • 评估49个模型发现其与人类水平仍有显著差距
  • 提供可微调语料库和无需训练的多智能体推理框架

现有对多模态大语言模型(MLLMs)空间智能的评估普遍碎片化且范围有限。本文提出SpatialScore,据我们所知,当前最全面、多样化的多模态空间智能评测基准。它覆盖多种视觉数据类型、输入模态与问答格式,包含约5000个手工验证样本,涵盖30种不同任务。基于该基准,我们系统评估了49个代表性MLLMs,揭示出当前模型在空间理解上存在持续性挑战,与人类水平存在显著差距。为提升模型能力,我们构建了包含33.1万条多模态问答样本的SpatialCorpus,支持空间推理任务的微调,并显著提升现有模型性能(如Qwen3-VL)。此外,为补充数据驱动路径,我们开发了SpatialAgent——一个配备12种专用空间感知工具的多智能体系统,支持Plan-Execute与ReAct推理,在无需额外训练的情况下实现空间推理能力的显著提升。大量实验与深入分析验证了所提基准、语料库与智能体框架的有效性。所有数据、代码与模型将向研究社区开源。

原文摘要 · Abstract (English)

Existing evaluations of multimodal large language models (MLLMs) on spatial intelligence are typically fragmented and limited in scope. In this work, we aim to conduct a holistic assessment of the spatial understanding capabilities of modern MLLMs and propose complementary data-driven and agent-based solutions. Specifically, we make the following contributions: (i) we introduce SpatialScore, to our knowledge, the most comprehensive and diverse benchmark for multimodal spatial intelligence to date. It covers multiple visual data types, input modalities, and question-answering formats, and contains approximately 5K manually verified samples spanning 30 distinct tasks; (ii) using SpatialScore, we extensively evaluate 49 representative MLLMs, revealing persistent challenges and a substantial gap between current models and human-level spatial intelligence; (iii) to advance model capabilities, we construct SpatialCorpus, a large-scale training resource with 331K multimodal QA samples that supports fine-tuning on spatial reasoning tasks and significantly improves the performance of existing models (e.g., Qwen3-VL); (iv) to complement this data-driven route with a training-free paradigm, we develop SpatialAgent, a multi-agent system equipped with 12 specialized spatial perception tools that supports both Plan-Execute and ReAct reasoning, enabling substantial gains in spatial reasoning without additional model training. Extensive experiments and in-depth analyses demonstrate the effectiveness of our benchmark, corpus, and agent framework. We expect these resources to serve as a solid foundation for advancing MLLMs toward human-level spatial intelligence. All data, code, and models will be released to the research community.

空间智能评测基准多模态推理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。