arXiv:2605.20165cs.CV2026-05

提出新评估框架,揭示视觉语言模型缺乏摄像机运动理解。

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models

论文配图:CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models
图 1 · 摘自论文原文
  • 设计空间叙事评分机制,要求模型描述场景与摄像机运动
  • 顶尖模型在新评测中性能大幅下降,暴露认知缺陷
  • 提出CaMo模型,实现空间理解能力的稳定表现

视觉语言模型(VLMs)在空间问答任务上表现优异,但其表现是否反映真正的空间智能仍不明确。我们发现现有空间VLMs缺乏基本的摄像机运动理解,这是空间认知的关键组成部分。为此,我们提出空间叙事评分(SNS),一种评估框架,要求VLM生成包含场景语义和摄像机运动的显式空间叙事,并通过冻结的代理大模型进行推理。在SNS下,最先进的空间VLMs性能显著下降,尽管其直接问答准确率很高。为弥补这一差距,我们引入了CaMo——一个基于摄像机运动的视觉语言模型,在SNS评估和直接空间问答准确率上均表现一致。结果表明,显式空间叙事外化对评估具备可迁移3D空间理解能力的VLM至关重要。代码、数据与模型已开源。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect genuine spatial intelligence. We show that existing spatial VLMs lack basic camera motion understanding, a key component of spatial cognition. We propose the Spatial Narrative Score (SNS), an evaluation framework that requires VLMs to generate explicit spatial narratives capturing both scene semantics and camera motion, followed by reasoning with a frozen proxy LLM. Under SNS, state-of-the-art spatial VLMs exhibit significant performance degradation despite high direct question answering accuracy. To address this gap, we introduce CaMo, a camera motion grounded VLM that achieves consistent performance across SNS evaluation and direct spatial question answering accuracy. Our results highlight the importance of explicit spatial narrative externalization for evaluating VLMs with transferable 3D spatial understanding. Our code, data, and model is available at https://github.com/hsiangwei0903/CaMo

视觉语言模型空间理解评估框架摄像机运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。