融合全局记忆与第一视角,提升长时程导航的决策能力
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
- 动态对齐全局记忆与局部视觉,增强空间推理
- 在物体导航任务中超越现有最先进方法
- 适合需要长期规划的智能体导航场景
近期大型语言模型(LLMs)和视觉-语言模型(VLMs)的发展使智能体在陌生环境中具备高效探索能力,能利用常识与空间推理。现有基于LLM的方法将全局记忆(如语义或拓扑地图)转化为语言描述以指导导航,虽提高效率并减少重复探索,但语言表示丢失几何信息,影响复杂环境中的空间推理。基于VLM的方法直接处理第一人称视觉输入以选择最优探索方向,但仅依赖局部视角导致部分可观测决策问题,在复杂场景中表现不佳。本文提出一种新型基于VLM的导航框架,通过自适应从全局记忆模块检索任务相关线索,并与代理的自我中心观测融合,实现全局上下文与局部感知的动态对齐,显著提升长时程任务中的空间推理与决策能力。实验表明,该方法在物体导航任务中优于现有最先进方法,提供更有效且可扩展的具身导航解决方案。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial reasoning for efficient exploration in unfamiliar environments. Existing LLM-based approaches convert global memory, such as semantic or topological maps, into language descriptions to guide navigation. While this improves efficiency and reduces redundant exploration, the loss of geometric information in language-based representations hinders spatial reasoning, especially in intricate environments. To address this, VLM-based approaches directly process ego-centric visual inputs to select optimal directions for exploration. However, relying solely on a first-person perspective makes navigation a partially observed decision-making problem, leading to suboptimal decisions in complex environments. In this paper, we present a novel vision-language model (VLM)-based navigation framework that addresses these challenges by adaptively retrieving task-relevant cues from a global memory module and integrating them with the agent's egocentric observations. By dynamically aligning global contextual information with local perception, our approach enhances spatial reasoning and decision-making in long-horizon tasks. Experimental results demonstrate that the proposed method surpasses previous state-of-the-art approaches in object navigation tasks, providing a more effective and scalable solution for embodied navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。