arXiv:2409.18794cs.ROcs.CV2024-09ICRA被引 77

用开源大模型实现零样本视觉语言导航,成本低且安全

Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs

  • 用时空思维链分解导航任务,提升理解与决策能力
  • 在仿真和真实环境均达闭源模型水平表现
  • 适合关注低成本、高安全性的AI导航研究者

视觉-语言导航(VLN)要求智能体根据文本指令在3D环境中导航。传统方法依赖领域特定数据集进行监督学习,而近期工作尝试使用GPT-4等闭源大模型实现零样本导航,但面临高昂的调用成本和潜在数据泄露风险。本文提出Open-Nav,首次探索在连续环境中使用开源大模型实现零样本VLN。该方法采用时空思维链(CoT)推理机制,将任务拆解为指令理解、进度估计与决策三个阶段,并通过细粒度物体与空间知识增强场景感知,提升大模型的导航推理能力。在仿真与真实环境中的大量实验表明,Open-Nav性能可媲美闭源大模型,同时显著降低使用成本与安全风险。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) tasks require an agent to follow textual instructions to navigate through 3D environments. Traditional approaches use supervised learning methods, relying heavily on domain-specific datasets to train VLN models. Recent methods try to utilize closed-source large language models (LLMs) like GPT-4 to solve VLN tasks in zero-shot manners, but face challenges related to expensive token costs and potential data breaches in real-world applications. In this work, we introduce Open-Nav, a novel study that explores open-source LLMs for zero-shot VLN in the continuous environment. Open-Nav employs a spatial-temporal chain-of-thought (CoT) reasoning approach to break down tasks into instruction comprehension, progress estimation, and decision-making. It enhances scene perceptions with fine-grained object and spatial knowledge to improve LLM's reasoning in navigation. Our extensive experiments in both simulated and real-world environments demonstrate that Open-Nav achieves competitive performance compared to using closed-source LLMs.

视觉导航开源模型零样本大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。