arXiv:2505.24257cs.CV2025-05EMNLP被引 7

测试视觉语言模型跨帧空间推理能力,发现其表现远低于人类。

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames

  • 设计新基准测试跨帧物体位置关系理解
  • 模型准确率随时间间隔增大从60%降至30%
  • 提供3D坐标可提升20%,凸显建模瓶颈

一个基于第一人称视角视频的具身AI助手需整合时间维度的空间线索——例如判断几秒前看到的物体A相对于稍后遇到的物体B的位置。我们提出Disjoint-3DQA,一个生成式问答基准,通过询问不在同一帧中共现的物体对关系来评估视觉语言模型(VLMs)的能力。我们测试了七种前沿VLMs,发现其性能比人类低28%,且随着时序间隔拉大,准确率从60%骤降至30%。分析显示,提供轨迹或鸟瞰图投影仅带来微小改善,而提供真实3D坐标则使性能提升20%。这揭示了多帧VLM在仅凭视觉信号构建并维持三维场景表示方面的核心瓶颈。Disjoint-3DQA因此为长时程空间推理设定了明确可测的挑战,旨在推动视觉、语言与具身AI交叉领域的研究。

原文摘要 · Abstract (English)

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We introduce Disjoint-3DQA , a generative QA benchmark that evaluates this ability of VLMs by posing questions about object pairs that are not co-visible in the same frame. We evaluated seven state-of-the-art VLMs and found that models lag behind human performance by 28%, with steeper declines in accuracy (60% to 30 %) as the temporal gap widens. Our analysis further reveals that providing trajectories or bird's-eye-view projections to VLMs results in only marginal improvements, whereas providing oracle 3D coordinates leads to a substantial 20% performance increase. This highlights a core bottleneck of multi-frame VLMs in constructing and maintaining 3D scene representations over time from visual signals. Disjoint-3DQA therefore sets a clear, measurable challenge for long-horizon spatial reasoning and aims to catalyze future research at the intersection of vision, language, and embodied AI.

空间推理视觉语言模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。