arXiv:2602.06974cs.ROcs.CV2026-02

用视觉相似性记忆实现无需地图的机器人导航

FeudalNav: A Simple Framework for Visual Navigation

  • 分层架构将导航拆解为多级决策,通过视觉相似性选择子目标
  • 在Habitat AI环境中表现媲美顶尖方法,训练与推理均不依赖里程计
  • 极少量人工干预即可显著提升导航成功率,适合交互式场景

机器人视觉导航受人类利用视觉线索和记忆导航能力启发,无需依赖详细地图。在未知、未测绘或无GPS的环境下,传统基于度量地图的方法失效,促使转向探索最少的学习型方法。本文提出一种分层框架,将导航决策过程分解为多个层次。方法通过一个简单且可迁移的航点选择网络学习选取子目标。关键组件是仅基于视觉相似性的潜在空间记忆模块,作为距离的代理。该替代图结构拓扑表示的方法足以完成导航任务,构建出紧凑、轻量、易于训练的导航器,可在新环境找到目标。我们在Habitat AI环境中展示了与多种前沿方法相当的结果,且训练与推理均未使用里程计。另一贡献在于利用框架的可解释性实现交互导航:我们探讨了达成所有试验成功的最小方向干预量。结果表明,即使极少的人工介入也能显著提升整体导航性能。

原文摘要 · Abstract (English)

Visual navigation for robotics is inspired by the human ability to navigate environments using visual cues and memory, eliminating the need for detailed maps. In unseen, unmapped, or GPS-denied settings, traditional metric map-based methods fall short, prompting a shift toward learning-based approaches with minimal exploration. In this work, we develop a hierarchical framework that decomposes the navigation decision-making process into multiple levels. Our method learns to select subgoals through a simple, transferable waypoint selection network. A key component of the approach is a latent-space memory module organized solely by visual similarity, as a proxy for distance. This alternative to graph-based topological representations proves sufficient for navigation tasks, providing a compact, light-weight, simple-to-train navigator that can find its way to the goal in novel locations. We show competitive results with a suite of SOTA methods in Habitat AI environments without using any odometry in training or inference. An additional contribution leverages the interpretablility of the framework for interactive navigation. We consider the question: how much direction intervention/interaction is needed to achieve success in all trials? We demonstrate that even minimal human involvement can significantly enhance overall navigation performance.

视觉导航分层策略交互式导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。