arXiv:2507.04047cs.CV2025-07ICCV被引 78

让智能体边走边理解3D场景,实现主动探索与视觉定位统一

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

  • 通过在线查询构建空间记忆,跳过复杂3D重建步骤
  • 联合优化物体定位与探索方向选择,提升导航效率14%以上
  • 支持语言、类别、图像等多种输入,适用于真实与模拟环境

具身场景理解不仅需要理解已观测的视觉空间信息,还需决定下一步在3D物理世界中探索何处。现有3D视觉-语言(3D-VL)模型主要聚焦于从静态3D重建(如网格和点云)中定位物体,但缺乏主动感知与探索能力。为此,我们提出Move to Understand(MTU3D),一个将主动感知与3D视觉-语言学习统一的框架,使智能体能有效探索并理解环境。核心创新包括:1)基于在线查询的表示学习,直接从RGB-D帧构建空间记忆,无需显式3D重建;2)统一的定位与探索目标,将未探索区域表示为前沿查询,联合优化物体定位与前沿选择;3)基于百万级多样化轨迹的端到端视觉-语言-探索预训练,覆盖仿真与真实世界的RGB-D序列。在多种具身导航与问答基准测试中,MTU3D在HM3D-OVON、GOAT-Bench、SG3D和A-EQA上的成功率分别优于最先进强化学习与模块化导航方法14%、23%、9%和2%。该模型具备多模态输入能力,支持类别、语言描述和参考图像,凸显了连接视觉定位与探索对具身智能的重要性。

原文摘要 · Abstract (English)

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus on grounding objects in static observations from 3D reconstruction, such as meshes and point clouds, but lack the ability to actively perceive and explore their environment. To address this limitation, we introduce \underline{\textbf{M}}ove \underline{\textbf{t}}o \underline{\textbf{U}}nderstand (\textbf{\model}), a unified framework that integrates active perception with \underline{\textbf{3D}} vision-language learning, enabling embodied agents to effectively explore and understand their environment. This is achieved by three key innovations: 1) Online query-based representation learning, enabling direct spatial memory construction from RGB-D frames, eliminating the need for explicit 3D reconstruction. 2) A unified objective for grounding and exploring, which represents unexplored locations as frontier queries and jointly optimizes object grounding and frontier selection. 3) End-to-end trajectory learning that combines \textbf{V}ision-\textbf{L}anguage-\textbf{E}xploration pre-training over a million diverse trajectories collected from both simulated and real-world RGB-D sequences. Extensive evaluations across various embodied navigation and question-answering benchmarks show that MTU3D outperforms state-of-the-art reinforcement learning and modular navigation approaches by 14\%, 23\%, 9\%, and 2\% in success rate on HM3D-OVON, GOAT-Bench, SG3D, and A-EQA, respectively. \model's versatility enables navigation using diverse input modalities, including categories, language descriptions, and reference images. These findings highlight the importance of bridging visual grounding and exploration for embodied intelligence.

具身智能3D导航视觉语言主动探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。