评测视觉语言模型在有约束空间中的结构化空间推理能力
Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces
- 构建基于真实3D结构的约束空间问答基准SSI-Bench
- 人类准确率91.6%,最强模型仅33.6%,差距显著
- 揭示当前模型在结构化空间推理中的根本缺陷
空间智能对视觉-语言模型至关重要,但多数场景中心基准评估的是无约束环境,单幅图像可能对应多种合理的3D解释。我们提出SSI-Bench,一个面向约束空间中结构中心空间推理(SCSR)的VQA基准。该基准基于复杂的现实世界3D结构,利用几何、拓扑和物理可行性等结构约束,使组件关系从视觉证据中更具确定性。基准包含1,000个排序类问题,涵盖几何与拓扑推理,正确排序需解析所有候选对象间的3D关系,对空间理解要求更高。其构建采用全人工中心化流程,耗时超400小时进行图像筛选、组件标注和问题设计。评估31个VLM发现与人类存在巨大差距:最佳开源模型准确率为22.2%,最强闭源模型达33.6%,而人类得分高达91.6%。进一步结果表明,思维链推理仅带来微弱提升,错误分析揭示当前模型在约束空间中空间理解的根本局限。项目页面:https://ssi-bench.github.io。
原文摘要 · Abstract (English)
Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a single image may admit multiple plausible 3D interpretations. We introduce SSI-Bench, a VQA benchmark for Structure-Centric Spatial Reasoning (SCSR) in constraint-governed spaces. Built from complex real-world 3D structures, it uses structural constraints from geometry, topology, and physical feasibility to make component relations more determinate from visual evidence. The benchmark contains 1,000 ranking questions spanning geometric and topological reasoning, where correct ordering requires resolving all candidate-wise 3D relations, imposing stronger demands on spatial understanding. It is created through a fully human-centered pipeline with over 400 researcher-hours of image curation, component annotation, and question design. Evaluating 31 VLMs reveals a large gap to humans: the best open-source model achieves 22.2% accuracy and the strongest closed-source model reaches 33.6%, while humans score 91.6%. Further results show that chain-of-thought reasoning brings only marginal gains, and error analysis reveals fundamental limitations in current models' spatial understanding within constraint-governed spaces. Project page: https://ssi-bench.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。