arXiv:2509.06266cs.CV2025-09被引 61

构建多视角场景空间推理新基准,提升视觉语言模型的三维理解能力

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes

  • 基于真实多视角户外数据构建评测基准Ego3D-Bench
  • 现有模型在空间问答上比人类低12%准确率,距离真实理解仍有差距
  • 提出可插拔框架Ego3D-VLM,显著提升距离估计与空间推理性能

当前视觉语言模型在理解三维空间关系方面仍存在明显局限。已有研究多基于单图或室内视频构建空间问答数据集,但现实中的机器人、自动驾驶等具身智能体通常依赖第一人称、多视角观测。为此,我们提出Ego3D-Bench,一个基于第一人称多视角户外数据的新型评测基准,用于评估视觉语言模型的空间推理能力。该基准包含超过8,600个问答对,由人工标注确保质量与多样性。我们对16个主流VLMs(如GPT-4o、Gemini1.5-Pro、InternVL3、Qwen2.5-VL)进行评测,结果表明模型表现与人类水平存在显著差距。为缩小这一差距,我们提出Ego3D-VLM——一种后训练框架,通过生成基于全局3D坐标估计的认知地图,使多选问答任务平均提升12%,绝对距离估计平均提升56%。该框架模块化设计,可适配任意现有VLM。Ego3D-Bench与Ego3D-VLM共同为实现真实多视角环境下的类人空间理解提供了重要工具。

原文摘要 · Abstract (English)

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos. However, real-world embodied AI agents such as robots and self-driving cars typically rely on ego-centric, multi-view observations. To this end, we introduce Ego3D-Bench, a new benchmark designed to evaluate the spatial reasoning abilities of VLMs using ego-centric, multi-view outdoor data. Ego3D-Bench comprises over 8,600 QA pairs, created with significant involvement from human annotators to ensure quality and diversity. We benchmark 16 SOTA VLMs, including GPT-4o, Gemini1.5-Pro, InternVL3, and Qwen2.5-VL. Our results reveal a notable performance gap between human level scores and VLM performance, highlighting that current VLMs still fall short of human level spatial understanding. To bridge this gap, we propose Ego3D-VLM, a post-training framework that enhances 3D spatial reasoning of VLMs. Ego3D-VLM generates cognitive map based on estimated global 3D coordinates, resulting in 12% average improvement on multi-choice QA and 56% average improvement on absolute distance estimation. Ego3D-VLM is modular and can be integrated with any existing VLM. Together, Ego3D-Bench and Ego3D-VLM offer valuable tools for advancing toward human level spatial understanding in real-world, multi-view environments.

空间推理视觉语言模型多视角具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。