arXiv:2605.20837cs.CVcs.AI2026-05被引 1

构建首个评估视觉语言模型建筑空间智能的基准,覆盖感知到配置五大维度。

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models

论文配图:ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models
图 1 · 摘自论文原文
  • 基于建筑学与认知科学设计五维评估框架,涵盖17个细粒度任务。
  • 构建3000组专家标注问答对,发现多数模型在空间变换与配置推理上仍落后人类。
  • 适合研究视觉语言模型空间理解能力或建筑智能的学者使用。

建筑空间智能是机器人导航、具身交互和三维场景理解与生成的基础能力。尽管现有研究已评估视觉语言模型(VLMs)在相对方向、距离比较和物体计数等基础空间技能上的表现,但这些任务仅涉及最底层的空间认知,忽视了布局理解、流线模式和功能分区等高层建筑空间认知。本文提出ArchSIBench,一个基于建筑学、认知科学与心理学视角的建筑空间智能基准。该基准包含感知、推理、导航、变换和配置五个核心维度,涵盖17个细粒度子任务。通过具有建筑背景的专家手工标注,构建了3000个问答对,实现对建筑空间智能的全面评估。基于此,我们评估了多种VLMs,发现大多数模型在建筑空间智能方面与人类基线存在显著差异,且不同能力维度间表现差异大。部分顶尖模型虽可接近无建筑训练的人类水平,但在空间变换与配置推理上仍明显落后于有建筑训练的人类评估者。我们认为ArchSIBench将为衡量与提升VLMs的建筑空间智能提供重要洞见与系统资源。数据集与代码已公开于https://huggingface.co/datasets/ArchSIBench/ArchSIBench。

原文摘要 · Abstract (English)

Architectural spatial intelligence, the ability to recognize and infer architectural space, is fundamental to tasks such as robot navigation, embodied interaction, and 3D scene understanding and generation. Although extensive research has evaluated the basic spatial skills of Vision-Language Models (VLMs) such as relative orientation, distance comparison, and object counting, these tasks cover only the most elementary levels of spatial cognition and largely overlook higher-level cognition of architectural space, including layout understanding, circulation patterns, and functional zoning. In this work, we present ArchSIBench, a Benchmark for Architectural Spatial Intelligence based on the perspectives from architecture, cognitive science, and psychology. ArchSIBench covers five core dimensions: perception, reasoning, navigation, transformation, and configuration, comprising 17 fine-grained subtasks. Through careful manual annotation by experts with architectural backgrounds, we construct 3,000 question-answer pairs to enable comprehensive evaluation of architectural spatial intelligence. Based on ArchSIBench, we evaluate various VLMs and find that the architectural spatial intelligence of most models shows significant differences from human baselines; additionally, models exhibit substantial variability across capability dimensions. Some state-of-the-art models can approach the level of human evaluators without architectural training. However, a clear gap remains compared to human evaluators with architectural training, particularly in spatial transformation and configuration reasoning. We believe that ArchSIBench will provide important insights and systematic resources for measuring and advancing the architectural spatial intelligence of VLMs. The dataset and code are available at https://huggingface.co/datasets/ArchSIBench/ArchSIBench.

建筑智能视觉语言模型空间推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。