让机器人用视觉语言模型实时探索未知环境并找目标
VL-Explore: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots
- 用CLIP模型实现视觉与语言的联合导航,边看边走
- 无需地图或先验知识,在真实场景中表现媲美有地图的算法
- 适合移动机器人在复杂环境中做零样本目标发现
视觉语言导航(VLN)已成为一种有前景的范式,使移动机器人能够在无需专门编程的情况下执行零样本推理。然而,现有系统通常将地图探索与路径规划分开,且由于环境信息有限(部分可见),探索效率较低。本文提出名为「VL-Explore」的新导航流程,可在未知环境中同时进行探索与目标发现,利用视觉语言模型CLIP的能力。该方法仅需单目视觉输入,无需预先构建地图或目标知识。为全面评估,我们设计并实现了一款名为「Open Rover」的无人地面车辆原型系统,用于通用VLN任务。将VL-Explore部署于Open Rover,测试其吞吐量、避障能力及轨迹性能。实验表明,VL-Explore在多种真实场景中持续优于传统地图遍历算法,并达到依赖先验地图与目标知识的路径规划方法的水平。特别地,该方法支持实时主动导航,无需预先采集候选图像或构建节点图,解决了现有VLN流程的关键局限。
原文摘要 · Abstract (English)
Vision-language navigation (VLN) has emerged as a promising paradigm, enabling mobile robots to perform zero-shot inference and execute tasks without specific pre-programming. However, current systems often separate map exploration and path planning, with exploration relying on inefficient algorithms due to limited (partially observed) environmental information. In this paper, we present a novel navigation pipeline named "VL-Explore" for simultaneous exploration and target discovery in unknown environments, leveraging the capabilities of a vision-language model named CLIP. Our approach requires only monocular vision and operates without any prior map or knowledge about the target. For comprehensive evaluations, we designed a functional prototype of a UGV (unmanned ground vehicle) system named "Open Rover", a customized platform for general-purpose VLN tasks. We integrated and deployed the VL-Explore pipeline on Open Rover to evaluate its throughput, obstacle avoidance capability, and trajectory performance across various real-world scenarios. Experimental results demonstrate that VL-Explore consistently outperforms traditional map-traversal algorithms and achieves performance comparable to path-planning methods that depend on prior map and target knowledge. Notably, VL-Explore offers real-time active navigation without requiring pre-captured candidate images or pre-built node graphs, addressing key limitations of existing VLN pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。