arXiv:2601.09954cs.CV2026-01被引 4

改进视觉语言模型的空间感知能力,提升对位置关系的理解。

The Spatial Blindspot of Vision-Language Models

  • 采用非拼贴式图像编码器和2D位置编码增强空间结构感知
  • 在多个基准测试中显著提升空间推理性能
  • 适合需要精准空间定位的机器人与具身智能应用

视觉语言模型(VLMs)虽发展迅速,但在捕捉空间关系方面仍存在盲点。当前多数VLM基于类似CLIP的图像编码器,训练时将图像展平为一维补丁序列,丢失了对空间推理至关重要的二维结构。本文认为,这种缺乏空间意识是VLM设计中的缺失环节,也是机器人和具身AI等需空间定位任务的瓶颈。为此,我们研究了(i)采用替代目标训练的图像编码器,以及(ii)2D位置编码。实验表明,这些架构选择可在多个基准上显著提升空间推理能力。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive language-image pretraining (CLIP) style image encoders. The training recipe often flattens images into 1D patch sequences, discarding the 2D structure necessary for spatial reasoning. We argue that this lack of spatial awareness is a missing dimension in VLM design and a bottleneck for applications requiring spatial grounding, such as robotics and embodied AI. To address this, we investigate (i) image encoders trained with alternative objectives and (ii) 2D positional encodings. Our experiments show that these architectural choices can lead to improved spatial reasoning on several benchmarks.

视觉语言模型空间推理位置编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。