用视觉大模型+规划器实现零样本物体导航,无需训练就能找新环境中的目标物。
SemNav: A Model-Based Planner for Zero-Shot Object Goal Navigation Using Vision-Foundation Models
- 结合视觉大模型与模型规划,通过探索前沿区域做长期决策。
- 在HM3D数据集上达到零样本导航最优的路径加权成功率。
- 适合需要快速适配新环境的机器人导航场景。
物体目标导航是具身智能中的基础任务,要求智能体在未知环境中定位指定目标物体。传统学习方法依赖大规模标注数据或强化学习中的大量环境交互,难以泛化到新环境且可扩展性差。为此,本文探索零样本设置,使智能体无需特定任务训练即可运行,提升可扩展性与适应性。近期视觉基础模型(VFMs)在视觉理解与推理方面表现强大,有助于智能体理解场景、识别相关区域并推断物体可能位置。本文提出一种零样本物体目标导航框架,融合VFMs的感知能力与具备长程决策能力的模型基规划器,通过前沿探索实现高效导航。我们在使用Habitat模拟器的HM3D数据集上评估该方法,结果表明其在零样本物体目标导航中取得了最优的路径加权成功率。
原文摘要 · Abstract (English)
Object goal navigation is a fundamental task in embodied AI, where an agent is instructed to locate a target object in an unexplored environment. Traditional learning-based methods rely heavily on large-scale annotated data or require extensive interaction with the environment in a reinforcement learning setting, often failing to generalize to novel environments and limiting scalability. To overcome these challenges, we explore a zero-shot setting where the agent operates without task-specific training, enabling more scalable and adaptable solution. Recent advances in Vision Foundation Models (VFMs) offer powerful capabilities for visual understanding and reasoning, making them ideal for agents to comprehend scenes, identify relevant regions, and infer the likely locations of objects. In this work, we present a zero-shot object goal navigation framework that integrates the perceptual strength of VFMs with a model-based planner that is capable of long-horizon decision making through frontier exploration. We evaluate our approach on the HM3D dataset using the Habitat simulator and demonstrate that our method achieves state-of-the-art performance in terms of success weighted by path length for zero-shot object goal navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。