评测视觉语言模型跨视角空间推理能力,发现其表现远低于人类。
CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming

- 构建卫星-街景配对数据集,支持多任务评估跨视角推理。
- 先进模型在视角剧烈变化时难以保持物体与布局一致性。
- 引入3D场景想象可显著提升模型跨视角推理能力,适合研究者参考。
人类能轻松进行跨视角场景推理,但视觉语言模型(VLMs)是否具备类似能力尚不明确。卫星-街景图像对因其复杂上下文和极端视角差异,成为理想测试场景。为此,我们提出CVSBench,一个大规模基准,用于通过卫星-街景配对评估跨视角空间推理能力。该基准包含3,297组跨视角图像、9,468个物体级标注和40,679个问答对,支持跨视角VQA、跨视角定位与视角识别等任务。大量实验表明,先进VLMs在剧烈视角变化下难以维持物体级和布局一致性。为逼近人类空间认知,我们探索两类方法:基于空间的推理机制与引入认知地图输入。结果表明,仅依赖语言的推理改进有限,而通过3D场景想象管道引入视觉空间想象力则显著提升性能。这凸显了显式视觉-空间表征对鲁棒空间认知的重要性。数据与代码已公开于https://huggingface.co/datasets/zlyzlyzly/CVSBench。
原文摘要 · Abstract (English)
Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities. Satellite-street scene pairs, with their complex contexts and extreme viewpoint variations, provide an ideal testbed. Motivated by this, we introduce CVSBench, a large-scale benchmark for evaluating cross-view spatial reasoning through satellite-street pairs. This benchmark supports multiple tasks, including cross-view VQA, cross-view grounding, and viewpoint identification. CVSBench comprises 3,297 cross-view image groups with 9,468 object-level annotations and 40,679 question-answer (QA) pairs, enabling systematic and controlled evaluation of cross-view spatial reasoning. Extensive evaluations reveal that advanced VLMs struggle to maintain object-level and layout consistency under drastic viewpoint changes. To bridge this gap towards human-like spatial cognition, we investigate two categories of approaches: spatially grounded reasoning and the incorporation of cognitive map inputs. Our findings demonstrate that language-only reasoning yields marginal improvements, while incorporating visual spatial imagination via a 3D scene imagination pipeline substantially improves cross-view reasoning. These results highlight the necessity of explicit visual-spatial representations for robust spatial cognition in VLMs. Our data and code are released at https://huggingface.co/datasets/zlyzlyzly/CVSBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。