arXiv:2509.02175cs.CVcs.AI2025-09被引 2

测试视觉语言模型的空间理解能力,发现只有顶级推理模型表现良好。

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

  • 构建新基准火箭科学,专注物体相对位置关系理解。
  • 主流VLM在空间关系任务上表现差,顶尖推理模型却显著领先。
  • 揭示空间推理是瓶颈,而非物体定位能力不足。

我们提出RocketScience,一个开源的对比性视觉语言模型(VLM)基准,用于测试空间关系理解能力。该基准由全新真实世界图像-文本对构成,主要涵盖相对空间关系及物体顺序。基准设计为人类易懂、当前VLM难以应对,经实证验证。结果表明,开源与前沿商业VLM普遍存在空间关系理解缺陷,而推理模型表现惊人优异。进一步解耦分析显示,基于思维链的模型性能受限于空间推理,而非物体定位能力。数据集采用CC-BY-4.0许可,评估代码已开源:https://github.com/nilshoehing/rocketscience。

原文摘要 · Abstract (English)

We propose RocketScience, an open-source contrastive VLM benchmark that tests for spatial relation understanding. It is comprised of entirely new real-world image-text pairs covering mostly relative spatial understanding and the order of objects. The benchmark is designed to be very easy for humans and hard for the current generation of VLMs, and this is empirically verified. Our results show a striking lack of spatial relation understanding in open source and frontier commercial VLMs and a surprisingly high performance of reasoning models. Additionally, we perform a disentanglement analysis to separate the contributions of object localization and spatial reasoning in chain-of-thought-based models and find that the performance on the benchmark is bottlenecked by spatial reasoning and not object localization capabilities. We release the dataset with a CC-BY-4.0 license and make the evaluation code available at: https://github.com/nilshoehing/rocketscience

空间理解视觉语言模型推理能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。