测试视觉语言模型在看不清时是否知道该放弃回答空间问题。
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?

- 设计新评测框架,模拟遮挡和视角误导两种真实场景挑战
- 模型在遮挡下准确率仅30%,视角模糊时低于10%
- 多数模型无法判断需补充哪个视角才能正确作答
空间推理是视觉语言模型在现实环境中部署的关键能力。然而,视觉观察本质上是对三维世界的有限表征:遮挡会使物体不可见,视角可能产生误导性几何线索。现有空间推理评测通常假设观测充分可靠,只关注模型能否给出正确答案,而忽视其是否能识别问题无法回答的情况,以及需要何种额外观测。本文构建了受控评测框架SpatialUncertain,引入两类观测挑战:(1) 遮挡,隐藏目标信息;(2) 视角模糊,产生误导性视觉线索。针对每种配置,设计在清晰观测下可答、但在挑战下应拒绝回答的空间问题。进一步评估模型能否识别哪些额外视角能解决视角模糊问题。对前沿开源与闭源VLM的测试显示两个一致失败模式:首先,模型易过度自信,在证据不全或误导时仍强行作答,遮挡下平均准确率约30%,视角模糊时低于10%;其次,即使有额外视角可用,部分模型在识别有效视角时表现接近随机水平。研究呼吁从单纯追求答案正确性转向评估模型是否懂得何时放弃及如何获取可靠证据。
原文摘要 · Abstract (English)
Spatial reasoning is a fundamental capability for vision-language models (VLMs) deployed in real-world environments. However, visual observations are inherently limited representations of a 3D world: occlusion can render objects invisible, and perspective can make geometric properties misleading. Despite this, existing spatial reasoning benchmarks typically assume that observations are sufficient and reliable, focusing on whether models produce correct answers rather than whether they recognize when a question cannot be answered and what additional observations would be needed. In this work, we challenge this assumption by constructing a controlled evaluation framework, SpatialUncertain, and introducing two types of observation challenges: (1) occlusion, which hides target information, and (2) perspective ambiguity, which produces misleading visual cues. For each configuration, we design spatial questions that are answerable under clean observations but require abstention under the introduced challenges. We further evaluate whether models can identify which additional viewpoints would resolve perspective ambiguity. Our results across a diverse set of frontier open- and closed-source VLMs reveal two consistent failure modes. First, models are prone to overconfident answering, attempting to solve spatial reasoning tasks even when visual evidence is incomplete or misleading, with average accuracy around 30\% under occlusion and below 10\% under perspective ambiguity. Second, even when additional views are available, some models perform near random chance in identifying which would provide reliable evidence. Together, our findings call for moving beyond answer correctness toward evaluating whether models know when to abstain and how to seek reliable evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。