用像素级相对3D图实现精准导航,无需全局几何一致
MASt3R-Nav: WayPixel Navigation in Relative 3D Maps

- 基于图像对间像素对应关系构建局部几何准确的连接图
- 通过稀疏化像素连接生成路径规划用的WayPixel代价图
- 在仿真和真实场景中均实现高精度自主导航,适合复杂环境
视觉导航能力高度依赖于对世界的表征。传统3D地图需全局几何一致性,而图像或物体相对的拓扑图虽简化了几何理解,却常限制导航仅能重复教学。本文提出一种像素相对连接的新地图表示,具备几何准确性但无需全局一致。受近期3D图像匹配进展启发,我们通过图像序列中成对图像的像素对应关系,在相对3D坐标系下构建地图。随后,通过近似与稀疏化图像内像素连接,实现全局路径规划,生成'WayPixel代价图',并训练以之为条件的控制器预测轨迹。实验表明,基于相对几何的密集像素代价图比图像或物体级别表征更适合作为控制预测的输入,显著提升导航性能,在四个类型的任务中通过仿真验证,并完成真实世界演示。
原文摘要 · Abstract (English)
Visual navigation ability is strongly tied to its underlying representation of the world. Unlike classical 3D maps that require globally-consistent geometry, image- or object-relative topological graphs almost entirely do away with geometric understanding. But, this comes at the cost of navigation capability, often limiting it to merely teach-and-repeat. In this work, we propose a novel map representation in the form of pixel-relative connectivity, which is geometrically accurate but does not require global geometric consistency. Inspired by recent progress in 3D grounded image matching, we construct a map from an image sequence through inter-image connectivity based on pixel correspondences in the relative 3D coordinate systems of individual image pairs. We then use this pixel-level graph to perform global path planning by approximating and sparsifying intra-image pixel connectivity. Through this, we derive a ''WayPixel Costmap'' representation and train a controller conditioned on it to predict a trajectory rollout. We show that this dense pixel-level costmap based on relative geometry is a more accurate conditioning variable for control prediction than its image- and object-level counterparts. This enables a highly capable navigation system, as validated on four types of navigation tasks in the simulator and through real world demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。