构建首个统一的空间推理评估基准,填补视觉语言模型在动态空间认知上的评测空白。
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- 基于认知科学分类法,将空间推理分为四类任务
- 生成超1.2万条可验证的训练数据,含559个评估样本
- 揭示当前模型在多视角多步骤推理上远低于人类水平
空间推理能力对视觉语言模型在机器人、增强现实、自动驾驶等领域的实际应用至关重要。然而,现有评测基准难以全面评估空间推理能力,尤其在人类空间认知中的内在-动态类型方面存在不足。本文提出统一基准Spatial-DISE,依据认知基础分类体系,将任务划分为四类:内在-静态、内在-动态、外在-静态、外在-动态。为解决数据稀缺问题,开发了可扩展的自动化生成管道,构建了包含559个评估VQA对的Spatial-DISE Bench和超过1.2万条训练VQA对的Spatial-DISE-12K数据集。对28个主流视觉语言模型的全面评估显示,当前模型在多步多视角空间推理上与人类能力存在显著且一致的差距。Spatial-DISE提供了可靠框架、宝贵数据集和未来研究方向,推动迈向类人空间智能。基准、数据集与代码将公开发布。
原文摘要 · Abstract (English)
Spatial reasoning ability is crucial for Vision Language Models (VLMs) to support real-world applications in diverse domains including robotics, augmented reality, and autonomous navigation. Unfortunately, existing benchmarks are inadequate in assessing spatial reasoning ability, especially the \emph{intrinsic-dynamic} spatial reasoning which is a fundamental aspect of human spatial cognition. In this paper, we propose a unified benchmark, \textbf{Spatial-DISE}, based on a cognitively grounded taxonomy that categorizes tasks into four fundamental quadrants: \textbf{I}ntrinsic-\textbf{S}tatic, Intrinsic-\textbf{D}ynamic, \textbf{E}xtrinsic-Static, and Extrinsic-Dynamic spatial reasoning. Moreover, to address the issue of data scarcity, we develop a scalable and automated pipeline to generate diverse and verifiable spatial reasoning questions, resulting in a new \textbf{Spatial-DISE} dataset that includes Spatial-DISE Bench (559 evaluation VQA pairs) and Spatial-DISE-12K (12K+ training VQA pairs). Our comprehensive evaluation across 28 state-of-the-art VLMs reveals that, current VLMs have a large and consistent gap to human competence, especially on multi-step multi-view spatial reasoning. Spatial-DISE offers a robust framework, valuable dataset, and clear direction for future research toward human-like spatial intelligence. Benchmark, dataset, and code will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。