arXiv:2605.23898cs.AI2026-05被引 1

测试视觉语言模型对空间数值的理解能力,发现其大多依赖表面线索,无法真正理解空间意义。

SPACENUM: Revisiting Spatial Numerical Understanding in VLMs

论文配图:SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
图 1 · 摘自论文原文
  • 构建双向任务评估视觉与数值间的空间映射能力
  • 多数模型在空间数值任务上表现接近随机猜测
  • 适合关注视觉语言模型泛化能力的开发者与研究者

视觉语言模型(VLMs)被越来越多应用于具身环境,需输出如动作幅度、空间坐标等数值。尽管这些数值看似合理,但其是否真正基于空间感知仍不明确。为此,本文提出SpaceNum框架,涵盖动态探索中的数值变化与静态布局中的空间推理两种场景,设计双向任务Num2Space和Space2Num,评估模型在视觉空间结构与语言数值表征之间的映射能力。我们系统检验当前VLMs在空间数值理解上的表现。在动态过渡与静态布局任务中,模型普遍无法将数值与空间意义有效关联,性能接近随机猜测。通过错误分析、推理轨迹分析与受控干预,发现现有模型严重依赖浅层空间线索,难以建立稳定的坐标感知表示,也无法从视觉观察中抽象出结构化的空间布局。进一步实验表明,显式推理仅带来微弱提升,而参数调优可部分改善空间数值理解并迁移到外部空间推理基准。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Although these numbers appear meaningful, it remains unclear whether these numerical outputs are genuinely grounded in spatial perception. Therefore, in this work, we revisit spatial numerical understanding through SpaceNum, a unified framework that captures two complementary settings: numbers as dynamic transitions during spatial exploration, and numbers as static layouts in spatial reasoning. We formulate two bidirectional tasks, Num2Space and Space2Num, to evaluate how well VLMs map between vision-side spatial structure and language-side numerical representations. We systematically study whether current VLMs truly understand numerical values in spatial settings. Across dynamic transitions and static layouts, we find that models largely fail to ground numbers in spatial meaning and often perform close to random guess. Through error analysis, reasoning trace analysis, and controlled interventions, we show that current VLMs rely heavily on shallow spatial cues, struggle to build stable coordinate-aware representations, and fail to abstract structured spatial layouts from visual observations. We further show that explicit reasoning provides only marginal gains, while tuning can partially improve spatial numerical understanding and transfer to external spatial reasoning benchmarks.

视觉语言模型空间理解数值推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。