arXiv:2506.03135cs.CVcs.AI2025-06被引 96

构建首个覆盖四类空间推理的综合评测基准,揭示视觉语言模型短板。

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

  • 基于认知心理学设计四类空间推理任务,覆盖动态、逻辑、交互与视角转换。
  • 构建8.4K高质量问答对,测试发现主流模型在复杂空间推理上表现不足。
  • 提出图结构提示与空间思维链新策略,提升模型空间推理能力。

空间推理是认知心理学的核心维度,也是当前视觉语言模型(VLMs)的瓶颈。尽管已有大量研究致力于评估或提升模型对基本空间关系(如左右、远近、物体计数)的理解,但这些任务仅涵盖最基础层次,且在最新模型中已趋于饱和。本文提出OmniSpatial,一个基于认知心理学的全面且具有挑战性的空间推理评测基准。该基准涵盖四大类别:动态推理、复杂空间逻辑、空间交互与视角转换,包含50个细粒度子类别。通过精心的人工标注,构建了超过8.4K个问答对。大量实验表明,无论是开源还是闭源的VLMs,在综合性空间推理任务上均存在显著局限。我们进一步探索两种增强策略:PointGraph(显式场景图提示)和SpatialCoT(新型视角思维链),以提升模型的空间推理性能。

原文摘要 · Abstract (English)

Spatial reasoning is a key aspect of cognitive psychology and remains a bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic spatial relations, such as distinguishing left from right, near from far, and object counting, these tasks cover only the most elementary layer of spatial reasoning and are largely approaching saturation in the latest reasoning models. In this work, we introduce OmniSpatial, a comprehensive and challenging benchmark for spatial reasoning, grounded in cognitive psychology. OmniSpatial covers four major categories: dynamic reasoning, complex spatial logic, spatial interaction, and perspective-taking, with 50 fine-grained subcategories. Through careful manual annotation, we construct over 8.4K question-answer pairs. Extensive experiments show that both open- and closed-source VLMs exhibit significant limitations in comprehensive spatial reasoning. We also explore two strategies-PointGraph (explicit scene graph cues) and SpatialCoT (novel-view chain-of-thought)-to bolster spatial reasoning.

空间推理视觉语言模型评测基准认知心理学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。