提出新型导航框架,让机器人零样本自主探索陌生环境。
DORAEMON: Decentralized Ontology-aware Reliable Agent with Enhanced Memory Oriented Navigation
- 分两路模拟人脑:一路处理空间语义融合,一路增强决策能力。
- 在三个数据集上成功率与路径效率均超越现有方法,提升显著。
- 无需地图预建或训练,适合真实家庭服务场景部署。
在陌生环境中实现自适应导航对家用服务机器人至关重要,但同时需要低层路径规划与高层场景理解,挑战巨大。尽管基于视觉语言模型(VLM)的零样本方法减少了对先验地图和特定场景训练数据的依赖,仍存在时空断续、记忆表示无结构、任务理解不足导致导航失败等问题。本文提出DORAEMON(去中心化语义感知可靠智能体,增强记忆导向导航),受人类认知启发,包含腹侧流与背侧流双路径架构。背侧流通过分层语义-空间融合与拓扑地图解决时空断续问题;腹侧流结合RAG-VLM与Policy-VLM提升决策能力。此外,设计Nav-Ensurance机制保障导航安全高效。在HM3D、MP3D和GOAT数据集上的评估显示,DORAEMON在成功率达(SR)与路径长度加权成功率(SPL)两项指标上均达到当前最优水平,显著优于现有方法。同时引入新评估指标AORI以更准确衡量导航智能性。大量实验表明,该方法可在无需地图构建或预训练的前提下实现零样本自主导航。
原文摘要 · Abstract (English)
Adaptive navigation in unfamiliar environments is crucial for household service robots but remains challenging due to the need for both low-level path planning and high-level scene understanding. While recent vision-language model (VLM) based zero-shot approaches reduce dependence on prior maps and scene-specific training data, they face significant limitations: spatiotemporal discontinuity from discrete observations, unstructured memory representations, and insufficient task understanding leading to navigation failures. We propose DORAEMON (Decentralized Ontology-aware Reliable Agent with Enhanced Memory Oriented Navigation), a novel cognitive-inspired framework consisting of Ventral and Dorsal Streams that mimics human navigation capabilities. The Dorsal Stream implements the Hierarchical Semantic-Spatial Fusion and Topology Map to handle spatiotemporal discontinuities, while the Ventral Stream combines RAG-VLM and Policy-VLM to improve decision-making. Our approach also develops Nav-Ensurance to ensure navigation safety and efficiency. We evaluate DORAEMON on the HM3D, MP3D, and GOAT datasets, where it achieves state-of-the-art performance on both success rate (SR) and success weighted by path length (SPL) metrics, significantly outperforming existing methods. We also introduce a new evaluation metric (AORI) to assess navigation intelligence better. Comprehensive experiments demonstrate DORAEMON's effectiveness in zero-shot autonomous navigation without requiring prior map building or pre-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。