arXiv:2603.07751cs.CVcs.CL2026-03中稿 · ICML被引 6

让视觉语言模型从2D图推3D空间,解决物体计数等基础难题。

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

  • 用正交投影分解场景,通过模拟与推理统一多视角认知
  • 在遮挡密集场景中物体计数准确率显著提升,视图一致性更强
  • 适合需要空间理解的多模态系统,如机器人导航、智能设计

当前大型语言模型已具备奥数级逻辑能力,但视觉语言模型在积木计数等基础空间任务上表现不佳。这种能力差异揭示了‘空间智能鸿沟’:模型无法从二维观察构建连贯的三维心理表征。诊断分析表明,瓶颈在于缺乏一致的视图空间接口,而非视觉特征不足或推理弱。为此,我们提出3ViewSense框架,将空间推理锚定于正交投影。借鉴工程认知原理,设计‘模拟与推理’机制,将复杂场景分解为标准正交投影以消除几何歧义。通过将第一人称感知与第三人称参考对齐,实现显式的心理旋转与重构。在空间推理基准测试中,该方法显著优于现有基线,在遮挡密集场景下的计数和视图一致性任务中均取得稳定提升。框架还增强了空间描述的稳定性和一致性,为多模态系统强化空间智能提供了可扩展路径。

原文摘要 · Abstract (English)

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,'' where models fail to construct coherent 3D mental representations from 2D observations. We uncover this gap via diagnostic analyses showing the bottleneck is a missing view-consistent spatial interface rather than insufficient visual features or weak reasoning. To bridge this, we introduce \textbf{3ViewSense}, a framework that grounds spatial reasoning in Orthographic Views. Drawing on engineering cognition, we propose a ``Simulate-and-Reason'' mechanism that decomposes complex scenes into canonical orthographic projections to resolve geometric ambiguities. By aligning egocentric perceptions with these allocentric references, our method facilitates explicit mental rotation and reconstruction. Empirical results on spatial reasoning benchmarks demonstrate that our method significantly outperforms existing baselines, with consistent gains on occlusion-heavy counting and view-consistent spatial reasoning. The framework also improves the stability and consistency of spatial descriptions, offering a scalable path toward stronger spatial intelligence in multimodal systems.~\footnote{https://github.com/Jasaxion/3ViewSense}

空间推理视觉语言模型正交投影多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。