提出首个面向物体中心空间推理的系统性评测基准。
Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- 基于可控合成数据集,评估模型对物体位置关系的理解能力。
- 发现检测器精确定位但缺乏关系推理,VLMs 生成流畅描述却难捕捉细节空间结构。
- 适合关注视觉模型空间理解能力的研究者和开发者参考。
空间理解是视觉基础模型的关键能力。尽管大型视觉模型或视觉语言模型(VLMs)在识别能力上取得进展,现有评测大多聚焦定位精度,而非模型是否真正理解场景中物体的排列与关联关系。这一差距至关重要:有效场景理解不仅需要识别物体,还需推理其相对位置、分组方式及深度关系。本文提出一个系统性的基础模型物体中心空间推理评测基准。采用受控合成数据集,评估了包括 GroundingDINO、Florence-2、OWLv2 等主流检测器,以及 InternVL、LLaVA、GPT-4o 等大型 VLMs 在三项任务上的表现:空间定位、空间推理与下游检索任务。结果揭示出稳定权衡:如 GroundingDINO 与 OWLv2 的检测器能提供精确边界框,但缺乏关系推理能力;而 SmolVLM 与 GPT-4o 等 VLMs 虽可生成粗略布局提示与流畅描述,但在细粒度空间上下文理解上表现不佳。研究凸显了定位与真实空间理解之间的鸿沟,呼吁社区发展更具空间感知能力的基础模型。
原文摘要 · Abstract (English)
Spatial understanding is a critical capability for vision foundation models. While recent advances in large vision models or vision-language models (VLMs) have expanded recognition capabilities, most benchmarks emphasize localization accuracy rather than whether models capture how objects are arranged and related within a scene. This gap is consequential; effective scene understanding requires not only identifying objects, but reasoning about their relative positions, groupings, and depth. In this paper, we present a systematic benchmark for object-centric spatial reasoning in foundation models. Using a controlled synthetic dataset, we evaluate state-of-the-art vision models (e.g., GroundingDINO, Florence-2, OWLv2) and large VLMs (e.g., InternVL, LLaVA, GPT-4o) across three tasks: spatial localization, spatial reasoning, and downstream retrieval tasks. We find a stable trade-off: detectors such as GroundingDINO and OWLv2 deliver precise boxes with limited relational reasoning, while VLMs like SmolVLM and GPT-4o provide coarse layout cues and fluent captions but struggle with fine-grained spatial context. Our study highlights the gap between localization and true spatial understanding, and pointing toward the need for spatially-aware foundation models in the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。