arXiv:2608.05747cs.CV2026-08被引 1

测试视觉语言模型从视频中构建全局空间认知的能力

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

论文配图:GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
图 1 · 摘自论文原文
  • 设计新基准GST-Bench,评估模型从连续视频推断全局空间关系
  • 最强模型仅得42.68分,远低于人类79.08分,暴露严重能力差距
  • 适合研究视频理解、空间推理与具身智能的学者参考

空间智能是具身智能体的基础,但现有基准多聚焦单帧或少数视角的局部空间感知,忽视了对连续长时程视觉流的全局空间认知。为此,我们提出全球时空基准(GST-Bench),一个面向视频理解中全局空间智能的VQA基准,包含由人工验证的6,790分钟合成视频生成的问题。该任务要求模型基于输入视频中未出现的新视角进行精确空间推理,并将第一人称观测映射到全局俯视图像。对22个先进视觉语言模型的全面评估显示,模型与人类之间存在显著差距:最强零样本模型得分仅42.68,远低于人类79.08。为探究原因,我们构建了GST-Bench-Local,发现即便在相同任务下模型具备良好的局部空间理解,仍无法将长时程观测整合为全局一致的场景表征。此外,我们还提供了用于全局空间推理的GST-Train数据集,以支持未来研究。

原文摘要 · Abstract (English)

Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-Spatial-Temporal Benchmark (GST-Bench), a VQA benchmark for global spatial intelligence in video understanding, comprising human-verified questions derived from 6,790 minutes of synthetically generated video. It requires models to perform accurate spatial inference from novel viewpoints unseen in the input video and to map egocentric observations onto global top-down images. A comprehensive evaluation of 22 state-of-the-art VLMs exposes a striking gap between models and humans: the strongest zero-shot model attains only 42.68, far below the human score of 79.08. To probe the cause of this gap, we construct GST-Bench-Local and find that models, despite strong local spatial understanding under the same task formulation, still fail to consolidate long-horizon observations into a globally consistent scene representation. We further provide GST-Train, a dataset for global spatial reasoning, as a complementary resource to facilitate future research on this challenge.

视频理解空间推理视觉语言模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。