用导航工具方向指令实现城市级视觉语言导航,无需密集地图或昂贵标注。
DA-Nav: Direction-Aware City-Scale Vision-Language Navigation

- 将导航转为基于视角图像的离散空间定位问题,通过思维链推理恢复轨迹。
- 在未见城市环境中达56.16%成功率,显著优于现有最先进方法。
- 无需微调即可适配四足与人形机器人,支持千米级闭环户外导航。
城市级室外导航目前严重依赖密集地图或高成本的导航监督。本文提出一种新范式,利用商业导航工具(如Google Maps)的方向性指令。为弥合商业指令与可执行导航动作之间的差距,并通过鲁棒轨迹恢复缓解长距离误差累积,我们提出DA-Nav——一种方向感知的视觉语言导航框架,将导航重新建模为在自参考2D图像平面上的离散空间定位问题。为实现轨迹恢复,DA-Nav采用包含偏差评估、动作预测和目标网格选择的思维链(CoT)推理过程。我们进一步构建ReDA数据集,提供方向感知指令与恢复轨迹,以增强空间定位并支持CoT恢复推理。在CARLA中的大量实验表明,DA-Nav在未见过的城市环境中实现了56.16%的高成功率,优于现有最先进方法,且具备更强的恢复能力。此外,无需微调,DA-Nav可无缝适配四足与人形机器人,在复杂真实环境中实现稳定的千米级闭环室外导航。
原文摘要 · Abstract (English)
City-scale outdoor navigation is currently hindered by the heavy reliance on dense maps or costly navigation supervision. In this work, we introduce a novel paradigm for leveraging directional instructions from commercial navigation tools (e.g., Google Maps). To bridge the gap between commercial instructions and executable navigation actions, while mitigating long-horizon error accumulation through robust trajectory recovery, we propose DA-Nav, a Direction-Aware vision-language Navigation framework that reformulates navigation as a discrete spatial grounding problem on the egocentric 2D image plane. To achieve trajectory recovery, DA-Nav employs a Chain-of-Thought (CoT) reasoning process encompassing deviation assessment, action prediction, and target grid selection. We further introduce ReDA, a dataset that provides direction-aware instructions and recovery trajectories to enhance spatial grounding and support CoT recovery reasoning. Extensive experiments in CARLA demonstrate that DA-Nav achieves a high success rate of 56.16% in unseen urban environments, outperforming existing State-of-The-Art (SoTA) methods while maintaining a substantially stronger recovery capability. Furthermore, without fine-tuning, DA-Nav seamlessly adapts to both quadruped and humanoid robots, enabling stable kilometer-scale closed-loop outdoor navigation in complex real world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。