arXiv:2410.02787cs.CVcs.AI2024-10被引 2

用视觉语言模型实现无需训练的智能导航,支持任意开放语言指令。

Navigation with VLM framework: Towards Going to Any Language

  • 以开源视觉语言模型为核心,直接理解自然语言目标进行导航
  • 在MP3D、HM3D等数据集上达成最优SPL指标,支持抽象与具体目标
  • 无需详细指令或环境先验,适合真实机器人在复杂室内场景应用

在开放环境中实现对任意语言目标的智能导航始终面临挑战。近年来,视觉语言模型(VLMs)展现出融合语言与视觉信息进行推理的强大能力。尽管已有诸多工作尝试利用VLM进行开放场景导航,但通常存在计算开销高、依赖物体中心方法或需依赖详细人类指令中的环境先验等问题。本文提出无需训练的导航框架NavVLM,利用开源视觉语言模型作为认知核心,使机器人能在开放场景中有效导航,即使面对抽象地点、动作或特定物体等人性化语言目标。NavVLM通过持续感知环境并提供探索引导,在仅提供简洁目标描述而无详细指令和环境先验的前提下实现智能导航。我们在仿真与真实世界中进行了验证:在仿真中,于Matterport 3D(MP3D)、Habitat Matterport 3D(HM3D)和Gibson丰富环境中,针对物体指定任务,其成功加权路径长度(SPL)达到当前最优水平;在真实机器人实验中,验证了该框架在真实室内场景下的有效性。

原文摘要 · Abstract (English)

Navigating towards fully open language goals and exploring open scenes in an intelligent way have always raised significant challenges. Recently, Vision Language Models (VLMs) have demonstrated remarkable capabilities to reason with both language and visual data. Although many works have focused on leveraging VLMs for navigation in open scenes, they often require high computational cost, rely on object-centric approaches, or depend on environmental priors in detailed human instructions. We introduce Navigation with VLM (NavVLM), a training-free framework that harnesses open-source VLMs to enable robots to navigate effectively, even for human-friendly language goal such as abstract places, actions, or specific objects in open scenes. NavVLM leverages the VLM as its cognitive core to perceive environmental information and constantly provides exploration guidance achieving intelligent navigation with only a neat target rather than a detailed instruction with environment prior. We evaluated and validated NavVLM in both simulation and real-world experiments. In simulation, our framework achieves state-of-the-art performance in Success weighted by Path Length (SPL) on object-specifc tasks in richly detailed environments from Matterport 3D (MP3D), Habitat Matterport 3D (HM3D) and Gibson. With navigation episode reported, NavVLM demonstrates the capabilities to navigate towards any open-set languages. In real-world validation, we validated our framework's effectiveness in real-world robot at indoor scene.

视觉语言模型机器人导航开放语言零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。