构建首个覆盖30类任务的多模态空间智能评测基准
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- 设计涵盖多种视觉数据与问答形式的综合评测集
- 评估49个模型发现其与人类水平仍有显著差距
- 提供可微调语料库和无需训练的多智能体推理框架
现有对多模态大语言模型(MLLMs)空间智能的评估普遍碎片化且范围有限。本文提出SpatialScore,据我们所知,当前最全面、多样化的多模态空间智能评测基准。它覆盖多种视觉数据类型、输入模态与问答格式,包含约5000个手工验证样本,涵盖30种不同任务。基于该基准,我们系统评估了49个代表性MLLMs,揭示出当前模型在空间理解上存在持续性挑战,与人类水平存在显著差距。为提升模型能力,我们构建了包含33.1万条多模态问答样本的SpatialCorpus,支持空间推理任务的微调,并显著提升现有模型性能(如Qwen3-VL)。此外,为补充数据驱动路径,我们开发了SpatialAgent——一个配备12种专用空间感知工具的多智能体系统,支持Plan-Execute与ReAct推理,在无需额外训练的情况下实现空间推理能力的显著提升。大量实验与深入分析验证了所提基准、语料库与智能体框架的有效性。所有数据、代码与模型将向研究社区开源。
原文摘要 · Abstract (English)
Existing evaluations of multimodal large language models (MLLMs) on spatial intelligence are typically fragmented and limited in scope. In this work, we aim to conduct a holistic assessment of the spatial understanding capabilities of modern MLLMs and propose complementary data-driven and agent-based solutions. Specifically, we make the following contributions: (i) we introduce SpatialScore, to our knowledge, the most comprehensive and diverse benchmark for multimodal spatial intelligence to date. It covers multiple visual data types, input modalities, and question-answering formats, and contains approximately 5K manually verified samples spanning 30 distinct tasks; (ii) using SpatialScore, we extensively evaluate 49 representative MLLMs, revealing persistent challenges and a substantial gap between current models and human-level spatial intelligence; (iii) to advance model capabilities, we construct SpatialCorpus, a large-scale training resource with 331K multimodal QA samples that supports fine-tuning on spatial reasoning tasks and significantly improves the performance of existing models (e.g., Qwen3-VL); (iv) to complement this data-driven route with a training-free paradigm, we develop SpatialAgent, a multi-agent system equipped with 12 specialized spatial perception tools that supports both Plan-Execute and ReAct reasoning, enabling substantial gains in spatial reasoning without additional model training. Extensive experiments and in-depth analyses demonstrate the effectiveness of our benchmark, corpus, and agent framework. We expect these resources to serve as a solid foundation for advancing MLLMs toward human-level spatial intelligence. All data, code, and models will be released to the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。