构建首个覆盖组合空间智能的多模态大模型评测基准
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- 定义10项基础空间能力,组合成8类复合空间任务
- 生成超5000个高质量问答对,覆盖811个真实室内场景
- 揭示计数能力不足严重制约模型组合推理能力
多模态大语言模型(MLLMs)在多种多模态任务中取得显著进展。为实现更高水平的空间智能,MLLMs需整合多项空间能力,即使面对简单常规任务亦然。然而,现有评测基准难以从原子到组合层面全面评估常见MLLMs的空间智能。为此,我们提出SpaCE-10,一个面向组合空间智能的综合性评测基准。在SpaCE-10中,我们定义了10项基础空间能力,并将其组合形成8类复合能力。基于此,我们设计了一种新型分层标注流程,生成高质量、多样化的问答对。经过超过150小时的人类专家投入,我们获得了超过5000个问答对,涵盖811个真实室内场景,支持点云输入与多选题等多种评测设置。我们在SpaCE-10上对常见MLLMs进行了广泛评估,发现即使最先进的模型与人类表现仍存在巨大差距。通过深入分析,我们得出若干重要结论,例如:当前MLLMs的计数能力缺陷是其组合空间能力受限的关键原因。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in various multimodal tasks. To pursue higher intelligence in space, MLLMs require integrating multiple spatial capabilities, even for handling simple and normal tasks. However, existing benchmarks struggle to comprehensively evaluate the spatial intelligence of common MLLMs from the atomic level to the compositional level. To fill this gap, we present SpaCE-10, a comprehensive benchmark for compositional spatial evaluations. In SpaCE-10, we define 10 atomic spatial capabilities, which are combined to form 8 compositional capabilities. Based on these definitions, we propose a novel hierarchical annotation pipeline to generate high-quality and diverse question-answer (QA) pairs. With over 150+ hours of human expert effort, we obtain over 5k QA pairs for 811 real indoor scenes in SpaCE-10, which covers various evaluation settings like point cloud input and multi-choice QA. We conduct an extensive evaluation of common MLLMs on SpaCE-10 and find that even the most advanced MLLM still lags behind humans by large margins. Through our careful study, we also draw several significant findings that benefit the MLLM community. For example, we reveal that the shortcoming of counting capability greatly limits the compositional spatial capabilities of existing MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。