arXiv:2608.06833cs.RO2026-08

无需时间顺序的图像导航,仅用普通照片就能让机器人找路。

Unordered Landmark Visual Navigation

论文配图:Unordered Landmark Visual Navigation
图 1 · 摘自论文原文
  • 纯视觉构建拓扑地图,通过几何验证和最大生成森林优化
  • 在无序图像下实现90%以上成功率,比现有方法提升超20%
  • 适合真实世界中使用随手拍照片训练的机器人系统

图像目标导航是具身智能的基础能力,但实际部署受限于强先验假设。现有方法多依赖时序视频流或深度、激光雷达等辅助传感器以维持空间一致性,这些顺序与多模态依赖严重制约可扩展性,尤其在使用众包或预录无序图像集时。移除时间先验后,现有方法面临严重感知混淆、噪声关联与灾难性建图失败。为此,我们提出无序地标视觉导航(ULVN),一种摆脱时间与里程计先验的统一RGB-only框架。ULVN通过集成建图、定位与规划系统性缓解误差累积:直接从非结构化图像构建鲁棒2D拓扑地图,采用校准几何验证与最大生成森林精炼;闭环执行中摒弃序列启发式,采用基于图的信念传播滤波器与熵自适应融合实现全局定位与动态子目标规划。仿真与真实世界实验表明,ULVN显著优于当前最优方法。

原文摘要 · Abstract (English)

Image-goal navigation is a fundamental capability for embodied AI, yet its practical deployment is strained by strong prior assumptions. Existing methods predominantly rely on temporally ordered video streams or auxiliary sensors (e.g., depth, LiDAR) to maintain spatial consistency. These sequential and multimodal dependencies severely restrict scalability, especially when deploying robots using crowd-sourced or pre-recorded unordered image collections. When temporal priors are removed, current methods struggle with severe perceptual aliasing, noisy associations, and catastrophic mapping failures. To address this underexplored challenge, we propose Unordered Landmark Visual Navigation (ULVN), a unified RGB-only framework free from temporal and odometric priors. ULVN systematically mitigates error accumulation by integrating mapping, localization, and planning. Specifically, it constructs a robust 2D topological map directly from unstructured images via calibrated geometric verification and maximum spanning forest refinement. For closed-loop execution, ULVN abandons sequential heuristics, utilizing a graph-based belief propagation filter with entropy-adaptive fusion for global localization and dynamic subgoal planning. Extensive experiments in simulation and real-world deployments demonstrate that ULVN significantly outperforms state-of-the-art methods.

视觉导航无序图像拓扑建图具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。