评测视觉语言模型的空间推理能力,聚焦视角变换下的物体关系理解。
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
- 以视角转换为切入点,设计分层诊断任务评估空间推理
- 43个模型均表现不佳,尤其在旋转和对称场景下错误率超50%
- 结果与人类认知难度高度相关,适合研究物理空间理解的学者
我们提出SpinBench,一个基于认知原理的视觉语言模型(VLM)空间推理诊断基准。该基准围绕视角转换的核心挑战——视角理解能力展开,要求模型在视点变化下推理场景和物体关系的变化。由于视角理解涉及跨视角识别、相对位置定位和心理模拟等多重认知能力,SpinBench设计了细粒度诊断类别,涵盖平移、旋转、物体相对姿态及视点变化,并按难度递进组织,从单物体简单任务逐步过渡到多物体复杂视角推理。我们评估了43个最先进的VLM模型(含开源与闭源)。结果揭示系统性弱点:显著的自我中心偏差、差的旋转理解能力,以及在对称与句法重构场景中的不一致性。规模分析显示性能随模型增大呈平滑提升并出现涌现能力。尽管人类在该任务中达到91.2%准确率,但任务难度(由人类反应时间衡量)与VLM准确率高度相关,表明SpinBench捕捉到了人类与模型共享的空间推理挑战。我们认为SpinBench为理解VLM在物理空间推理中的局限提供了关键洞见。
原文摘要 · Abstract (English)
We present SpinBench, a cognitively grounded diagnostic benchmark for evaluating spatial reasoning in vision language models (VLMs). SpinBench is designed around the core challenge of spatial reasoning: perspective taking, the ability to reason about how scenes and object relations change under viewpoint transformation. Since perspective taking requires multiple cognitive capabilities, such as recognizing objects across views, relative positions grounding, and mentally simulating transformations, SpinBench introduces a set of fine-grained diagnostic categories. Our categories target translation, rotation, object relative pose, and viewpoint change, and are progressively structured so that single-object simpler tasks scaffold toward the most demanding multi-object perspective-taking setting. We evaluate 43 state-of-the-art VLMs, both proprietary and open source. Results reveal systematic weaknesses: strong egocentric bias, poor rotational understanding, and inconsistencies under symmetrical and syntactic reformulations. Scaling analysis shows both smooth improvements and emergent capabilities. While human subjects achieve high accuracy (91.2\%), task difficulty as measured by human response time shows strong correlation with VLM accuracy, indicating that SpinBench captures spatial reasoning challenges shared across humans and VLMs. We believe SpinBench provides critical insights into spatial reasoning in VLMs and highlights key gaps in their ability to reason about physical space. Our website can be found at https://spinbench25.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。