揭示视觉语言模型的空间认知盲区,推动更类人空间推理能力发展
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation
- 构建分层评估框架,从基础到高阶逐步测试模型空间推理能力
- 发现主流模型在距离、视角和物理逻辑推理上存在显著缺陷
- 适合研究视觉语言模型、空间认知与人机交互的学者参考
当前视觉语言模型虽能理解基本空间线索和简单方向(如左、右、前、后),但在实现类人空间理解与真实应用所需的多维空间推理方面仍存在困难。为弥补这一差距,我们提出SPHERE(Spatial Perception and Hierarchical Evaluation of REasoning)——一个基于新的人工标注数据集的分层评估框架。SPHERE系统性地在复杂度递增的层次上测试模型,涵盖基础技能、多技能整合以及结合空间、视觉与逻辑理解的高阶推理。对先进模型的基准测试显示,其在距离与邻近关系判断、自我中心与非自我中心视角理解,以及物理情境中的空间逻辑应用方面存在明显不足。这些发现揭示了现有模型的关键盲点,凸显了发展更先进空间推理技术的必要性,推动视觉语言模型向更贴近人类空间认知的方向演进。SPHERE基准测试数据集已开源:https://github.com/zwenyu/SPHERE-VLM。
原文摘要 · Abstract (English)
Current vision-language models may grasp basic spatial cues and simple directions (e.g. left, right, front, back), but struggle with the multi-dimensional spatial reasoning necessary for human-like understanding and real-world applications. To address this gap, we develop SPHERE (Spatial Perception and Hierarchical Evaluation of REasoning), a hierarchical evaluation framework supported by a new human-annotated dataset. SPHERE systematically probes models across increasing levels of complexity, from fundamental skills to multi-skill integration and high-level reasoning that combines spatial, visual, and logical understanding. Benchmark evaluation of state-of-the-art models reveals significant deficiencies, especially in reasoning about distance and proximity, understanding both egocentric and allocentric perspectives, and applying spatial logic in physical contexts. These findings expose critical blind spots in existing models and underscore the need for more advanced spatial reasoning techniques, driving the development of vision-language models that align more closely with human spatial cognition. The SPHERE benchmark is available at https://github.com/zwenyu/SPHERE-VLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。