用博弈论方法让机器人理解复杂语言指令,零样本导航更准更快
VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation
- 构建3D物体中心地图,融合视觉语言特征与环境重建
- 在HM3D数据集上超越现有方法,零样本任务成功率提升12.7%
- 适合需要理解复杂描述的机器人导航场景
遵循人类指令在陌生环境中探索并寻找指定目标,是移动服务机器人的关键能力。以往的物体目标导航研究多仅依赖单一模态输入,难以充分考虑包含详细属性和空间关系的语言描述。为此,我们提出VLN-Game,一种新型零样本视觉目标导航框架,可有效处理物体名称和描述性语言目标。该方法通过整合预训练视觉语言特征与环境3D重建,构建3D物体中心空间地图,并识别最值得探索的潜在目标区域。采用博弈论视觉语言模型判断哪个目标最匹配给定语言描述。在Habitat-Matterport 3D(HM3D)数据集上的实验表明,该框架在物体目标导航和基于语言的导航任务中均达到当前最优性能。此外,证明了VLN-Game可直接部署于真实机器人。结果凸显了使用紧凑视觉语言模型结合博弈论方法,显著提升机器人决策能力的潜力。
原文摘要 · Abstract (English)
Following human instructions to explore and search for a specified target in an unfamiliar environment is a crucial skill for mobile service robots. Most of the previous works on object goal navigation have typically focused on a single input modality as the target, which may lead to limited consideration of language descriptions containing detailed attributes and spatial relationships. To address this limitation, we propose VLN-Game, a novel zero-shot framework for visual target navigation that can process object names and descriptive language targets effectively. To be more precise, our approach constructs a 3D object-centric spatial map by integrating pre-trained visual-language features with a 3D reconstruction of the physical environment. Then, the framework identifies the most promising areas to explore in search of potential target candidates. A game-theoretic vision language model is employed to determine which target best matches the given language description. Experiments conducted on the Habitat-Matterport 3D (HM3D) dataset demonstrate that the proposed framework achieves state-of-the-art performance in both object goal navigation and language-based navigation tasks. Moreover, we show that VLN-Game can be easily deployed on real-world robots. The success of VLN-Game highlights the promising potential of using game-theoretic methods with compact vision-language models to advance decision-making capabilities in robotic systems. The supplementary video and code can be accessed via the following link: https://sites.google.com/view/vln-game.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。