让视觉语言模型通过想象未来视角,实现无需地图的智能导航。
ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- 用图像想象构建未来视点,将导航转化为选最佳视角任务。
- 在开放词汇导航基准上超越现有方法,无需预先地图。
- 适合研究视觉导航与多模态决策的学者和工程师。
视觉导航是家庭服务机器人完成日常长时任务的关键能力,尤其体现在物品搜寻上。当前许多方法利用大语言模型(LLM)进行常识推理以提升探索效率,但其规划过程局限于文本,难以仅通过文本表达空间占据与几何布局,而这对于合理导航决策至关重要。本文旨在释放视觉语言模型(VLM)的空间感知与规划能力,探究仅依赖车载摄像头捕获的RGB/RGB-D流输入,VLM能否高效完成无地图的视觉导航任务。为此,我们提出ImagineNav框架,通过生成关键机器人视角下的未来观察图像,将复杂的导航规划简化为VLM可处理的最佳视角选择问题。我们引入Where2Imagine模块,通过蒸馏使候选视角生成符合人类导航习惯。最终,采用现成的点目标导航策略引导机器人到达目标视角。在具有挑战性的开放词汇物体导航基准上的实验表明,所提系统显著优于现有方法。
原文摘要 · Abstract (English)
Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve exploration efficiency. However, the planning process of LLMs is limited within texts and it is difficult to represent the spatial occupancy and geometry layout only by texts. Both are important for making rational navigation decisions. In this work, we seek to unleash the spatial perception and planning ability of Vision-Language Models (VLMs), and explore whether the VLM, with only on-board camera captured RGB/RGB-D stream inputs, can efficiently finish the visual navigation tasks in a mapless manner. We achieve this by developing the imagination-powered navigation framework ImagineNav, which imagines the future observation images at valuable robot views and translates the complex navigation planning process into a rather simple best-view image selection problem for VLM. To generate appropriate candidate robot views for imagination, we introduce the Where2Imagine module, which is distilled to align with human navigation habits. Finally, to reach the VLM preferred views, an off-the-shelf point-goal navigation policy is utilized. Empirical experiments on the challenging open-vocabulary object navigation benchmarks demonstrates the superiority of our proposed system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。