空间视觉语言模型看似一致,实则错判距离,缺乏真实视觉证据支持。
Consistent Yet Wrong: Evidence Insensitivity in Spatial Vision-Language Models

- 构建多视角评估框架ViewDiag,检验模型对不同视角下物体距离的判断
- 多数模型在80个场景中表现高度一致但准确率低,错误仍稳定输出
- 揭示模型依赖先验而非视觉证据,适合研究空间认知与鲁棒性的人参考
空间推理是机器人、自主系统和具身智能的基础,但当前视觉语言模型(VLMs)在度量距离查询上仍不可靠。通常认为跨视角预测一致即代表几何理解,我们验证发现相反:主流VLMs即使答案错误也常保持视点不变,表明预测与具体视觉证据耦合薄弱。本文提出ViewDiag,基于Hypersim、ScanNet和KITTI360构建的可控多视角评估协议,包含176个物体对轨迹,覆盖80个场景,每条轨迹有2至10个视角。该协议从度量准确性、分布集中度及内部坍塌三方面评估模型,其中内部坍塌通过潜在特征探针检测。结果显示,各模型普遍呈现高稳定性与高误差并存,处于强一致性但低准确性的区域。这挑战了以跨视角一致性作为几何理解代理的常用做法。我们证明,稳定预测可能源于先验驱动的坍塌,而非证据敏感推理。ViewDiag提供了一个控制性基准与诊断框架,用于检验空间VLM是否不仅准确,更真正依赖视觉证据。
原文摘要 · Abstract (English)
Spatial reasoning is fundamental to robotics, autonomy, and embodied AI, yet modern vision-language models (VLMs) remain unreliable on metric distance queries. A common assumption is that consistent predictions across viewpoints reflect geometric grounding. We test this assumption and find the opposite: leading VLMs often produce view-invariant and consistent answers even when those answers are incorrect, indicating weak coupling between predictions and viewpoint-specific visual evidence. We introduce \textbf{ViewDiag}, a controlled multi-view evaluation protocol built from Hypersim, ScanNet, and KITTI360, comprising 176 object-pair tracks across 80 scenes with 2--10 views per track. The protocol evaluates models along three axes: metric accuracy, distributional concentration, and internal collapse, the last of which is assessed using a latent feature probe. Across diverse models, we observe a consistent pattern of high prediction stability paired with substantial error, clustering in a regime characterized by strong consistency but low accuracy. \noindent These results challenge the common use of cross-view consistency as a proxy for geometric understanding. Instead, we show that stable predictions may reflect prior-driven collapse rather than evidence-sensitive reasoning. ViewDiag provides a controlled benchmark and diagnostic framework for evaluating whether spatial VLMs are not only accurate, but also meaningfully coupled to visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。