让多模态大模型直接用俯视图地图导航,实现零样本物体寻物。
TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
- 基于俯视图地图进行空间推理,避免语言转换导致的信息丢失。
- 在MP3D和HM3D数据集上达到领先性能,导航成功率显著提升。
- 适合研究智能体导航、空间认知与多模态大模型应用的读者。
零样本物体导航(ZSON)任务要求具身智能体在陌生环境中寻找从未见过的物体。此类目标导向的探索高度依赖对环境空间信息的感知、理解与推理能力。然而,现有基于大语言模型的方法将视觉观测转化为语言描述并在语言空间中推理,导致空间信息损失。本文提出TopV-Nav,一种基于多模态大模型(MLLM)的方法,直接在包含丰富空间信息的俯视图地图上进行推理。为充分挖掘MLLM在俯视视角下的空间推理潜力,我们提出自适应视觉提示生成(AVPG)方法,动态构建语义丰富的俯视图地图,使智能体能直接利用地图中的空间信息进行深度推理。此外,设计动态地图缩放(DMS)机制,按需调整地图缩放层级,增强局部细粒度推理能力。同时引入潜在目标驱动(PTD)机制,预测并利用目标位置,促进全局性、类人探索行为。在MP3D和HM3D数据集上的实验表明,TopV-Nav具有显著优势。
原文摘要 · Abstract (English)
The Zero-Shot Object Navigation (ZSON) task requires embodied agents to find a previously unseen object by navigating in unfamiliar environments. Such a goal-oriented exploration heavily relies on the ability to perceive, understand, and reason based on the spatial information of the environment. However, current LLM-based approaches convert visual observations to language descriptions and reason in the linguistic space, leading to the loss of spatial information. In this paper, we introduce TopV-Nav, an MLLM-based method that directly reasons on the top-view map with sufficient spatial information. To fully unlock the MLLM's spatial reasoning potential in top-view perspective, we propose the Adaptive Visual Prompt Generation (AVPG) method to adaptively construct semantically-rich top-view map. It enables the agent to directly utilize spatial information contained in the top-view map to conduct thorough reasoning. Besides, we design a Dynamic Map Scaling (DMS) mechanism to dynamically zoom top-view map at preferred scales, enhancing local fine-grained reasoning. Additionally, we devise a Potential Target Driven (PTD) mechanism to predict and to utilize target locations, facilitating global and human-like exploration. Experiments on MP3D and HM3D datasets demonstrate the superiority of our TopV-Nav.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。