测试视觉语言模型从上下文推断说话者视角的能力
Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models

- 构建新数据集POVBench,分离三种视角理解任务
- 现有模型在方向性语言定位上仍表现不佳
- 显式分解空间推理能提升定位准确率
在机器人等具身任务中,理解语言指令常需从说话者所处场景视角推理空间关系。人类可借助共享环境知识、活动背景和常识推断视角。尽管近期视觉语言模型(VLMs)展现出空间推理能力,但其能否通过上下文线索推断说话者视角,并据此理解情境化空间关系尚不明确。我们称此能力为上下文观察者定位(contextual observer grounding)。为此,我们构建了视角基准测试(POVBench),包含3D场景与查询,将推断、陈述和已知三种形式的观察者定位解耦。给定多视角观测和自然语言句子,模型需根据情境化空间线索定位未见或描述不清的目标。在多种先进VLMs中,即使明确给出观察者定位,基于方向性语言的目标定位仍具挑战。我们发现,显式分解观察者相对的空间推理可提升目标定位性能。
原文摘要 · Abstract (English)
Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated perspective. Humans infer such perspectives from shared environmental knowledge, activity context, and commonsense. Recent vision-language models (VLMs) appear capable of spatial reasoning, but their ability to infer a speaker's viewpoint from contextual cues and interpret situated spatial relations from that viewpoint remains unclear. We call this capability contextual observer grounding. To study this capability, we construct the Point-of-View Benchmark (POVBench), a dataset of 3D scenes and queries that disentangles Inferred, Stated, and Given forms of observer grounding in natural embodied communication. Given multi-view observations and a natural-language sentence, models must localize unseen or underspecified targets from situated spatial and contextual cues. Across state-of-the-art VLMs, localizing targets from directional language remains challenging, even when observer grounding is made explicit. We find that explicit breakdowns of observer-relative spatial reasoning improve target localization. Our project page is available at https://mimo-owl.github.io/POVBench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。