arXiv:2409.02522cs.AIcs.RO2024-09被引 5

用大模型构建会思考的导航智能体,让机器人像人一样理解语言并走迷宫。

Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments

  • 基于大模型构建认知地图,融合时间、空间与语义信息
  • 通过预测路径点优化探索效率,提升导航成功率至78.3%
  • 模拟人类反思机制,支持持续学习和动态调整策略

视觉-语言导航在连续环境(VLN-CE)中代表了具身智能的前沿挑战,要求智能体仅凭自然语言指令在无界三维空间中自由导航。该任务对多模态理解、空间推理和决策能力提出高要求。为此,我们提出基于大语言模型(LLMs)的生成式智能体Cog-GA,专为VLN-CE设计。Cog-GA采用双路径策略模拟人类认知过程:首先构建融合时空语义元素的认知地图,帮助LLM建立空间记忆;其次通过路径点预测机制,战略性优化探索轨迹以提升导航效率。每个路径点配备双通道场景描述,将环境线索分为‘是什么’与‘在哪里’两类,类似人脑处理方式,增强注意力聚焦,精准提取导航所需空间信息。此外,反射机制捕获过往导航反馈,支持持续学习与自适应重规划。在VLN-CE基准上的大量评估验证了Cog-GA的顶尖性能,其表现出类人的导航行为,显著推动了战略性与可解释性智能体的发展。

原文摘要 · Abstract (English)

Vision Language Navigation in Continuous Environments (VLN-CE) represents a frontier in embodied AI, demanding agents to navigate freely in unbounded 3D spaces solely guided by natural language instructions. This task introduces distinct challenges in multimodal comprehension, spatial reasoning, and decision-making. To address these challenges, we introduce Cog-GA, a generative agent founded on large language models (LLMs) tailored for VLN-CE tasks. Cog-GA employs a dual-pronged strategy to emulate human-like cognitive processes. Firstly, it constructs a cognitive map, integrating temporal, spatial, and semantic elements, thereby facilitating the development of spatial memory within LLMs. Secondly, Cog-GA employs a predictive mechanism for waypoints, strategically optimizing the exploration trajectory to maximize navigational efficiency. Each waypoint is accompanied by a dual-channel scene description, categorizing environmental cues into 'what' and 'where' streams as the brain. This segregation enhances the agent's attentional focus, enabling it to discern pertinent spatial information for navigation. A reflective mechanism complements these strategies by capturing feedback from prior navigation experiences, facilitating continual learning and adaptive replanning. Extensive evaluations conducted on VLN-CE benchmarks validate Cog-GA's state-of-the-art performance and ability to simulate human-like navigation behaviors. This research significantly contributes to the development of strategic and interpretable VLN-CE agents.

视觉导航大模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。