用方向思维替代盲目探索,让机器人导航更高效准确
DRIVE-Nav: Directional Reasoning, Inspection, and Verification for Efficient Open-Vocabulary Navigation
- 以持久方向为导航核心,减少重复检查和无效路径
- 在HM3D-OVON上达50.2%成功率,比前人提升1.9个百分点
- 适合需要高效自主导航的机器人与真实场景部署
开放词汇物体导航(OVON)要求智能体在未知环境中定位语言描述的目标。现有零样本方法依赖不完整观测下的前沿候选推理,拓扑感知方法虽减少冗余但仍有全景检查开销和重复判断问题。本文提出DRIVE-Nav,一种以持久方向为中心的结构化框架。通过更全面地检查遇到的方向,并将后续决策限制在前向240度视域内的相关方向,有效减少重复访问,提升路径效率。该框架从加权快速行进法(FMM)路径中提取并追踪方向性候选,维护代表性视觉信息用于语义检查,并结合视觉-语言引导的提示增强与跨帧验证,提高定位可靠性。在HM3D-OVON、HM3Dv1、HM3Dv2和MP3D上的实验表明,DRIVE-Nav表现优异且具持续效率优势。在HM3D-OVON上实现50.2%成功率(SR)和32.6%路径相似度(SPL),分别优于此前最佳方法1.9%和5.6%。同时在HM3Dv1、HM3Dv2和MP3D上取得最优SPL,且成功迁移到物理人形机器人,在真实场景中验证有效性。
原文摘要 · Abstract (English)
Open-Vocabulary Object Navigation (OVON) requires an embodied agent to locate a language-specified target in unknown environments. Many zero-shot methods rely on frontier-candidate reasoning under incomplete observations, while topology-aware methods reduce candidate redundancy but may still introduce panoramic inspection overhead and repeated reconsideration. We present DRIVE-Nav, a structured framework that organizes exploration around persistent directions rather than raw frontiers. By inspecting encountered directions more completely and restricting subsequent decisions to still-relevant directions within a forward 240-degree view range, DRIVE-Nav reduces redundant revisits and improves path efficiency. The framework extracts and tracks directional candidates from weighted Fast Marching Method (FMM) paths, maintains representative views for semantic inspection, and combines vision-language-guided prompt enrichment with cross-frame verification to improve grounding reliability. Experiments on HM3D-OVON, HM3Dv1, HM3Dv2, and MP3D demonstrate strong overall performance and consistent efficiency gains. On HM3D-OVON, DRIVE-Nav achieves 50.2% SR and 32.6% SPL, improving the previous best method by 1.9% SR and 5.6% SPL. It also delivers the best SPL on HM3Dv1, HM3Dv2, and MP3D and transfers to a physical humanoid robot. Real-world deployment also demonstrates its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。