arXiv:2606.03988cs.AI2026-06

通过想象感知令牌提升视觉语言模型的空间推理能力

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

论文配图:Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
图 1 · 摘自论文原文
  • 引入想象感知令牌,模拟未观测视角下的视觉认知
  • 在多视角计数任务上准确率提升3.4%,优于纯文本思维链
  • 适合需要空间推断的视觉理解场景,如机器人导航

视觉语言模型(VLMs)在多项任务中表现优异,但在关键信息不可见时仍难以进行空间推理。这类问题需依赖想象感知:推断从未观测视角可见内容、追踪遮挡空间中的路径,或整合部分观察形成连贯的空间表征。本文提出想象感知令牌(IPT),作为中间感知表示,外化模型在不同空间配置下应感知的内容,同时与观测输入保持一致。为评估该能力,构建了三个任务——视角切换(PET)、路径追踪(PT)和多视角计数(MVC),并建立约2万条带真实想象标注的数据集及评测基准。以BAGEL为骨干模型,IPT监督在多个任务中持续提升空间推理性能,甚至在不生成图像的情况下超越纯文本思维链训练。在MVC任务中准确率提升3.4%,在PT任务中达到与强闭源模型相当的水平。进一步发现,IPT与仅标签监督结合可带来额外增益,而文本思维链可能因模态错配显著降低性能。总体而言,IPT为未观测空间结构的推理提供了合理监督信号,在提升泛化能力的同时生成可解释的中间表示。

原文摘要 · Abstract (English)

Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require imaginative perception: inferring what would be seen from an unseen viewpoint, tracing paths through occluded spaces, or integrating partial observations into a coherent spatial representation. We introduce Imaginative Perception Tokens (IPT), intermediate perceptual representations that externalize what a VLM would perceive under alternative spatial configurations while remaining consistent with the observed input. To study this capability, we formulate three tasks, Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC), and construct datasets of approximately 20K examples with ground truth imaginations, answers, and evaluation benchmarks. Using the unified VLM BAGEL as the backbone, IPT supervision consistently improves spatial reasoning and often outperforms textual chain of thought training, even without generating images at inference time. On MVC, IPT improves accuracy by 3.4% and achieves competitive performance with strong closed-source models on PT. We further find that combining IPT and label-only supervision yields additional gains, whereas textual chain of thought can substantially degrade performance, suggesting a modality mismatch when spatial computation is forced through language. Overall, IPT provides a principled supervision signal for reasoning about unobserved spatial structure, improving generalization while producing interpretable intermediate representations.

空间推理视觉语言模型感知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。