融合深度与视觉语言信息,提升机器人导航效率与泛化能力。
EffiNav: Fusing Depth and Vision-Language for Efficient Object Goal Navigation

- 结合深度图与视觉语言模型,动态规划探索路径。
- 在HM3D和OVON上达成更高成功率与路径效率。
- 适用于仿真与真实机器人,适应性强且不易重复探索。
在未知环境中定位目标物体是自主智能体的核心能力,广泛应用于搜救与野外机器人任务。对象目标导航(ObjNav)要求智能体高效抵达目标,其路径效率反映了探索策略的优劣。现有训练型模型存在泛化不足,非训练框架则易出现重复探索或往返移动。本文提出EffiNav,融合深度信息与视觉语言模型,在Habitat Matterport 3D(HM3D)与开放词汇对象目标导航(OVON)两个主流仿真基准上评估,并在真实机器人上验证有效性。通过大规模仿真失败分析,仅需微小修改即可扩展至带记忆的GOAT-BENCH任务。在成功率(SR)与路径长度加权成功率(SPL)两项标准下,EffiNav表现优于或相当最新基线,体现其高效、鲁棒且实用的特点,尤其在不同数据集间展现出更均衡的泛化能力。
原文摘要 · Abstract (English)
To locate a target object while exploring the unknown environment is a fundamental capability for autonomous agents, with applications ranging from search-and-rescue to field robots. A simplified version of such task is Object Goal Navigation (ObjNav). In ObjNav, successful arrival at the target object provides a basic measure of performance; however, the efficiency of the navigation trajectory is equally important, as it indicates how intelligently the agent explores and how much time remains for subsequent tasks. In unknown environments, the key to efficient navigation lies in deciding where to explore next. While many prior works aim to address this core challenge and achieved promising performance in certain settings, recent training-based models and non-training frameworks still suffer from generalization and efficiency issues respectively, which in the worst cases can lead to excessive exploration of already-visited areas or redundant back-and-forth motion. We evaluate EffiNav on two widely used simulation benchmarks Habitat Matterport 3D (HM3D) and Open-Vocabulary Object goal Navigation (OVON), and further validate its effectiveness on physical robots in real-world settings. We conduct failure analysis on massive simulation episodes. With minimal modification, we also extend EffiNav to a memory-augmented ObjNav task on the GOAT-BENCH dataset, demonstrating its adaptability beyond standard ObjNav settings. Across two standard metrics--Success Rate (SR) and Success weighted by Path Length (SPL), EffiNav matches or outperforms recent baselines, reflecting its efficiency, robustness, and practical applicability. Recognizing the different emphases of the two datasets, the performances reveals this framework is more balanced and generalizable for efficient ObjNav.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。