arXiv:2511.19119cs.CV2025-11中稿 · ECCV被引 4

构建首个面向单目图像的开放词汇空间推理数据集

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

  • 提出MonoSR数据集,支持多种场景与问题类型
  • 发现现有视觉语言模型在单目空间推理上表现有限
  • 揭示辅助信息对提升推理效果的关键作用

空间推理(SR)是指从二维输入中推断三维空间信息的能力,在具身智能和自动驾驶等真实场景中至关重要。然而,现有研究主要聚焦于室内环境,且通常依赖多视角观测,限制了其在室外场景和单目图像这一最常见的现实设置中的泛化能力。本文提出MonoSR,一个大规模单目空间推理数据集,覆盖室内、室外及以物体为中心的多样化场景,并支持多种问题类型。MonoSR为开放世界单目空间推理提供了新路径。除了引入数据集,我们还评估了先进的视觉-语言模型,揭示其在此挑战性任务上的局限性。进一步分析了辅助信息对单目空间推理的重要性,为未来模型设计提供实用指导。这些贡献共同奠定了在真实开放世界环境中推进单目空间推理的基础。

原文摘要 · Abstract (English)

Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and autonomous driving. However, existing research primarily focuses on indoor environments and typically relies on multi-view observations, which limits their generalizability to outdoor scenarios and constrains their applicability to monocular images, the most common real-world setting. In this work, we propose MonoSR, a large-scale monocular spatial reasoning dataset that spans diverse scenarios including indoor, outdoor, and object-centric settings, and supports multiple question types. MonoSR provides a path toward open-world monocular spatial reasoning. Beyond introducing the dataset, we evaluate advanced vision-language models to reveal their limitations on this challenging task. We further analyze whether auxiliary information is crucial for monocular spatial reasoning and offer practical guidance for designing future models. These contributions collectively establish a foundation for advancing monocular spatial reasoning in real-world, open-world environments.

空间推理单目图像视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。