arXiv:2601.22228cs.CVcs.AI2026-01被引 1

视觉语言模型在跨视角定位任务上表现差,暴露出空间推理短板。

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation

  • 将相对相机位姿估计转为文本分类任务,构建真实数据集验证
  • 顶级模型仅达0.66准确率,远低于人类的0.91和几何方法的0.99
  • 尤其在旋转与深度变化上严重失准,适合诊断多视图推理缺陷

本文研究视觉语言模型(VLM)是否能从图像对中解决相对相机位姿估计(RCPE)问题,作为多视角空间推理的直接测试。将RCPE转化为离散文本分类任务,构建了基于真实RGB-D帧的VRRPI-Bench,以及用于隔离单一运动自由度的VRRPI-Diag。人类(0.91)和专用几何流水线如LoFTR(0.99)能可靠完成任务,但最佳VLM仅达0.66,多数接近随机水平。分析表明,该差距并非基础空间能力缺失:强VLM在单图基准上接近天花板,但跨视图推理时性能骤降。其在源-目标反转下一致性最低仅59.7%,且在简化单自由度设置中仍表现薄弱,尤其在光学轴运动如俯仰和深度平移上(GPT-5在滚动上仅0.46)。这些失败揭示出关键缺失能力:跨视图对应、视角一致推理及投影相机运动理解,使RCPE成为提升VLM多视图空间推理能力的精准诊断工具。

原文摘要 · Abstract (English)

We study whether vision-language models (VLMs) can solve relative camera pose estimation (RCPE) from image pairs, a direct test of multi-view spatial reasoning. We cast RCPE as a discrete verbal classification task and introduce \texttt{VRRPI-Bench}, built from real RGB-D frames with object-centric camera motion, and \texttt{VRRPI-Diag}, which isolates individual motion degrees of freedom. Humans (0.91) and specialized geometric pipelines such as LoFTR (0.99) solve the task reliably, yet the best VLM reaches only 0.66 and most others remain near random. Our analyses show that this gap is not basic spatial competence: strong VLMs are near ceiling on single-image benchmarks, but most remain near random once reasoning must span views. They are unstable under source-target reversal (best 59.7\% consistency) and remain weak even in simplified single-DoF settings, especially on optical-axis motions such as roll and depth translation (GPT-5: 0.46 on roll). These failures are useful: they localize concrete missing capabilities, namely cross-view correspondence, view-consistent reasoning, and projective camera-motion understanding, making RCPE a targeted diagnostic for improving multi-view spatial reasoning in VLMs.

视觉语言模型空间推理相机位姿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。