构建新基准评估视觉模型的空间推理能力
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
- 设计可控场景的合成图像与空间标注,测试模型相对位置识别能力
- 多模型对比显示空间推理能力差异显著,现有模型普遍不足
- 适合关注视觉模型空间认知、具身智能研究者使用
视觉基础模型(如 DINO、CLIP)在图像语义理解上表现优异,但在空间推理方面能力有限,制约其在具身系统中的应用。尽管近期研究引入深度估计等3D任务进行训练,模型在其他空间任务上表现仍不一致,引发对其是否具备真正空间意识的质疑。为此,我们提出空间关系识别任务(SpaRRTa)基准,评估模型识别图像中物体相对位置的能力。与侧重精确度量预测(如表面法线估计)的传统3D目标不同,SpaRRTa考察更基础的人类级空间理解能力。该基准生成任意数量的逼真图像,涵盖多样场景和可控制的物体布局,并提供开放获取的空间标注。对一系列先进视觉基础模型的评估揭示了其空间推理能力的显著差异。通过分析,我们深入探讨了支撑或阻碍现代视觉基础模型空间意识的机制。期望 SpaRRTa 能为未来空间感知视觉模型的发展提供指导。
原文摘要 · Abstract (English)
Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems. As a result, recent work incorporates some 3D tasks (such as depth estimation) into VFM training. However, VFM performance remains inconsistent across other spatial tasks, raising the question of whether these models truly have spatial awareness or overfit to specific 3D objectives. To address this question, we introduce the Spatial Relation Recognition Task (SpaRRTa) benchmark, which evaluates the ability of VFMs to identify relative positions of objects in the image. Unlike traditional 3D objectives that focus on precise metric prediction (e.g., surface normal estimation), SpaRRTa probes a fundamental capability underpinning more advanced forms of human-like spatial understanding. SpaRRTa generates an arbitrary number of photorealistic images with diverse scenes and fully controllable object arrangements, along with freely accessible spatial annotations. Evaluating a range of state-of-the-art VFMs, we reveal significant disparities between their spatial reasoning abilities. Through our analysis, we provide insights into the mechanisms that support or hinder spatial awareness in modern VFMs. We hope that SpaRRTa will serve as a useful tool for guiding the development of future spatially aware visual models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。