arXiv:2603.06985cs.CV2026-03被引 2

让视觉语言模型更准地理解单目图像中的空间关系。

Perception-Aware Multimodal Spatial Reasoning from Monocular Images

  • 用视觉参考令牌统一表示物体,实现视觉与文本联合推理。
  • 在SURDS基准上超越此前方法,单/多物体任务均显著提升。
  • 适合需要精准空间感知的自动驾驶场景研究者。

从单目图像进行空间推理对自动驾驶至关重要,但现有视觉-语言模型仍难以应对尺度变化大和物体外观模糊的情况。本文提出一种感知感知的多模态空间推理框架,赋予视觉语言模型显式的以物体为中心的定位能力。不依赖文本框输出,每个被指代物体通过其空间范围内所有视觉参考令牌(VRTs)表示,使视觉证据与文本推理可在统一标记空间中协同处理。为强化跨模态交互,构建了对齐视觉与文本推理信号的多模态思维链(MM-CoT)数据集,并引入确定性排序策略,使对无序VRT集合的监督与VLM的自回归预测完全兼容。仅通过标准监督微调,该方法在SURDS基准上取得显著提升,超越此前使用强化学习后训练的方法,在单物体与多物体任务中表现均大幅领先。结果表明,精确感知与多模态推理相互促进,共同构成复杂单目驾驶场景下鲁棒空间理解的关键。

原文摘要 · Abstract (English)

Spatial reasoning from monocular images is essential for autonomous driving, yet current Vision-Language Models (VLMs) still struggle with fine-grained geometric perception, particularly under large scale variation and ambiguous object appearance. We propose a simple yet effective perception-aware multimodal reasoning framework that equips VLMs with explicit object-centric grounding ability. Instead of relying on textual bounding-box outputs, each referred object is represented using all Visual Reference Tokens (VRTs) within its spatial extent, enabling visual evidence and textual reasoning to be processed jointly in a unified token space. To further strengthen cross-modal interaction, we construct a Multimodal Chain-of-Thought (MM-CoT) dataset that injects aligned visual and textual reasoning signals. A deterministic ordering strategy is introduced to make supervision over inherently unordered VRT sets fully compatible with the VLM's autoregressive next-token prediction. With only standard supervised fine-tuning, our method achieves substantial improvements on the SURDS benchmark, outperforming previous approaches - including those using RL-based post-training - by a large margin across both single-object and multi-object tasks. These results demonstrate that accurate perception and multimodal reasoning are mutually reinforcing, and together form the key to robust spatial understanding in challenging monocular driving scenarios.

空间推理视觉语言模型自动驾驶多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。