arXiv:2511.21471cs.AI2025-11被引 29

构建首个分层空间认知评测框架,揭示大模型在空间推理中的能力短板。

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

  • 提出五级递进的空间认知框架,覆盖从观察到规划的完整链条。
  • 在15个任务上测试海量多模态大模型,发现其感知强但规划弱。
  • 适合关注空间智能评估与模型局限性的研究者参考。

空间认知是实现真实世界多模态智能的基础,使模型能有效与物理环境互动。尽管多模态大语言模型(MLLMs)已取得显著进展,现有评测常将空间认知简化为单一维度指标,无法捕捉其层级结构与内在依赖关系。为此,我们提出一种分层空间认知框架,将空间智能分解为五个逐步复杂的层级:从基础观测到高层规划。基于此分类,我们构建了SpatialBench——一个大规模、细粒度的基准,涵盖15个与认知层级对齐的任务。为实现异构任务间的统一评估,我们引入面向高阶能力的量化指标,可靠衡量模型的整体空间推理能力。大量实验表明,模型在感知锚定方面表现良好,但在符号推理、因果推断和规划方面仍受限。额外的人类测试显示,人类具有选择性、目标导向的抽象能力,而MLLMs则过度关注表面细节,缺乏连贯的空间意图。本工作首次系统建立用于评估MLLMs中分层空间认知的框架,为未来空间智能系统奠定基础。

原文摘要 · Abstract (English)

Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides, existing benchmarks often oversimplify spatial cognition, reducing it to a single-dimensional metric, which fails to capture the hierarchical structure and interdependence of spatial abilities. To address this gap, we propose a hierarchical spatial cognition framework that decomposes spatial intelligence into five progressively complex levels from basic observation to high-level planning. Building upon this taxonomy, we construct SpatialBench, a large-scale, fine-grained benchmark covering 15 tasks aligned with these cognitive levels. To provide a unified evaluation across heterogeneous tasks, we further introduce a high-level capability-oriented metric that reliably assesses a model's overall spatial reasoning ability. Extensive experiments over massive MLLMs reveal distinct performance stratification across cognitive levels: models exhibit strong perceptual grounding yet remain limited in symbolic reasoning, causal inference, and planning. Additional human tests demonstrate that humans perform selective, goal-directed abstraction, while MLLMs tend to over-attend to surface details without coherent spatial intent. Our work establishes the first systematic framework for measuring hierarchical spatial cognition in MLLMs, laying the foundation for future spatially intelligent systems.

空间认知多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。