arXiv:2605.06317cs.CVcs.AI2026-05

提出一步全局规划方法,让导航更高效准确

NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps

论文配图:NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps
图 1 · 摘自论文原文
  • 将导航转为单步全局路径规划,直接输出密集路径概率
  • 在新数据集上达最优性能,规划速度比基线快80倍
  • 适合需要快速高精度导航的机器人应用

现有视觉语言导航方法多采用渐进式、逐步推理模式,易累积误差且效率低。尽管近期方法尝试利用预构建环境地图,但常依赖逐步更新记忆图谱或评分离散路径候选,限制了连续空间推理并引入离散瓶颈。本文提出自上而下视觉语言导航(TD-VLN),将导航重构为基于预建俯视地图的一步全局路径规划问题,并构建了新数据集R2R-TopDown。为此提出NavOne统一框架,通过端到端单次前向传播直接预测多模态地图上的密集路径概率。NavOne包含俯视地图融合模块实现多模态联合表征,并扩展注意力残差机制实现空间感知深度融合。在R2R-TopDown上的大量实验表明,NavOne在基于地图的VLN方法中达到最先进水平,规划阶段速度相较现有地图基线提升8倍,相比原生姿态方法快80倍,实现高效全局导航。

原文摘要 · Abstract (English)

Existing Vision-Language Navigation (VLN) methods typically adopt an egocentric, step-by-step paradigm, which struggles with error accumulation and limits efficiency. While recent approaches attempt to leverage pre-built environment maps, they often rely on incrementally updating memory graphs or scoring discrete path proposals, which restricts continuous spatial reasoning and creates discrete bottlenecks. We propose Top-Down VLN (TD-VLN), reformulating navigation as a one-step global path planning problem on pre-built top-down maps, supported by our newly constructed R2R-TopDown dataset. To solve this, we introduce NavOne, a unified framework that directly predicts dense path probabilities over multi-modal maps in a single end-to-end forward pass. NavOne features a Top-Down Map Fuser for joint multi-modal map representation, and extends Attention Residuals for spatial-aware depth mixing. Extensive experiments on R2R-TopDown show that NavOne achieves state-of-the-art performance among map-based VLN methods, with a planning-stage speedup of 8x over existing map-based baselines and 80x over egocentric methods, enabling highly efficient global navigation.

视觉语言导航路径规划地图建模机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。