用单目视觉构建3D空间地图,让机器人导航更安全
MVP-Nav: Multi-layer Value Map Planner Navigator

- 通过2D语义图生成3D边界框,重建物理占据信息
- 多层价值图融合语义与几何,在无深度情况下实现精准规划
- 零样本导航性能领先,适合无深度传感器的智能体
仅使用RGB图像进行零样本物体目标导航(ZSON)对具身智能体构成根本挑战,因缺乏显式深度信息导致严重的物理不确定性与语义-物理错位。现有方法或依赖高层语义推理而无几何基础,或学习端到端策略但缺乏显式物理约束,常产生语义合理却物理不安全的行为。本文提出MVP-Nav,一种物理感知的纯RGB导航框架,将感知、规划与控制与真实3D世界对齐。MVP-Nav利用3D基础模型将2D语义实例投影为3D方向边界框,从单目观测中重建显式物理占据,形成全局空间语义表征。为统一高层语义推理与底层物理约束,引入多层价值图(MVM),将语义优先级与重构几何整合至共享代价空间,实现物理基础的几何规划。在零样本物体导航基准上的大量实验表明,MVP-Nav显著优于现有无深度方法,达到当前最佳性能,验证了结构化物理先验可有效弥补主动深度传感器缺失的不足。
原文摘要 · Abstract (English)
Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment. Existing approaches either rely on high-level semantic reasoning without geometric grounding or learn end-to-end policies that lack explicit physical constraints, often resulting in semantically plausible but physically unsafe behaviors. In this paper, we propose MVP-Nav, a physical-aware RGB-only navigation framework that aligns perception, planning, and control with the real 3D world. MVP-Nav reconstructs explicit physical occupancy from monocular observations by leveraging 3D foundation models to project 2D semantic instances into 3D oriented bounding boxes, forming a global spatial semantic representation. To unify high-level semantic reasoning and low-level physical constraints, we introduce a Multi-layer Value Map (MVM) that integrates semantic priorities and reconstructed geometry into a shared cost space, enabling physically grounded geometric planning. Extensive experiments on zero-shot object navigation benchmarks demonstrate that MVP-Nav significantly outperforms existing depth-free methods, achieving state-of-the-art performance and validating that structured physical priors can effectively compensate for the absence of active depth sensors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。