arXiv:2606.09268cs.RO2026-06

仅用单目摄像头实现精准定位与障碍物度量感知,兼顾低成本与安全性。

VGP-Nav: Metric-Aware Visual Geometric Perception for Robot Navigation

论文配图:VGP-Nav: Metric-Aware Visual Geometric Perception for Robot Navigation
图 1 · 摘自论文原文
  • 利用地面几何约束锚定视觉尺度,解决单目系统固有尺度模糊问题
  • 在线实时生成具度量一致性的障碍物地图,支持后续路径规划
  • 无需多传感器融合,适合低成本机器人部署,实测表现稳定

可靠的机器人导航需要无缝集成精确的全局定位与密集、度量一致的障碍物感知。传统方法依赖相机提供丰富视觉特征用于定位,同时使用激光雷达等主动传感器获取直接度量数据,但多传感器配置需复杂的时空标定且增加部署开销。尽管纯视觉方案成本低、可扩展,现有单目系统难以同时实现高效、全局一致的定位与密集度量一致的几何感知。为此,本文提出VGP-Nav——一种仅依赖单目RGB输入的统一框架,实现度量感知下的视觉几何感知。核心思想是将基于定位的视觉几何与来自地面平面几何的物理意义尺度约束相锚定,从而为单目感知提供可靠的度量参考。VGP-Nav在运行中在线解决单目尺度模糊问题,生成基于定位的、具度量一致性的障碍物表示,可直接用于下游规划。大量实验表明其在多样环境中的强泛化能力,并成功部署于真实移动机器人,验证了该方法在可扩展、低成本、安全自主导航中的实用性。

原文摘要 · Abstract (English)

Reliable robotic navigation necessitates the seamless integration of accurate global localization and dense, metric-consistent obstacle perception. A common strategy to achieve these capabilities involves integrating diverse sensing modalities: cameras offer rich visual features for localization, while active sensors like LiDAR provide direct metric measurements. However, such multi-sensor configurations necessitate complex spatial-temporal calibration and increase deployment overhead. Although vision-only approaches offer a low-cost and scalable alternative, existing monocular visual systems typically struggle to simultaneously achieve efficient, globally consistent localization and dense, metric-consistent geometric perception. To bridge this gap, we propose \textbf{VGP-Nav}, a unified framework for \textit{Metric-Aware Visual Geometric Perception} that relies solely on monocular RGB input to jointly support metric localization and obstacle perception. Our key insight is to anchor localization-grounded visual geometry to physically meaningful scale constraints derived from ground-plane geometry, thereby providing a reliable metric reference for monocular perception. VGP-Nav resolves monocular scale ambiguity online and produces localization-grounded, metric obstacle representations that are directly applicable to downstream planning. Extensive experiments demonstrate strong generalization across diverse environments and successful deployment on real mobile robots, highlighting the practicality of our approach for scalable, low-cost, and safe autonomous navigation.

机器人导航单目视觉度量感知几何推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。