用记忆图与空间场结合,让机器人长期导航更懂语义、会探索。
CGFM-Nav: Cognitive Graph-Field Memory for Semantic-Guided Lifelong Multimodal Embodied Navigation

- 构建记忆图与语义前沿场,兼顾显式记忆与连续探索引导
- 在GOAT-Bench上成功率提升至63.0%,SPL达39.6%
- 适合需要长期多模态导航的智能体系统研究者
视觉语言导航(VLN)要求智能体在持续探索未知区域的同时,对累积观测进行推理。然而,现有环境表征难以同时支持显式语义记忆与连续探索引导。为此,我们提出认知图-场记忆(CGFM),一种持久的多模态场景表示,将对象、空间关系与视觉观测组织为多模态场景图,实现跨任务的长程目标检索与推理。当无法识别可靠目标时,基于图的证据被投影到目标条件化的语义前沿场中,引导探索向语义上有潜力的区域推进。基于CGFM,我们提出CGFM-Nav,一个基于基础模型的终身多模态导航框架,将任务相关的子图选择、视觉语言模型推理与验证反馈整合进闭环决策流程。在相同Qwen3-VL-8B骨干网络下,初步实验显示,在GOAT-Bench上整体成功率从53.2%提升至63.0%,SPL从30.0%提升至39.6%,证明了显式语义记忆与语义引导探索结合的有效性。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) requires agents to reason over accumulated observations while continuously exploring unseen regions. However, existing environment representations often struggle to jointly support explicit semantic memory and continuous exploration guidance. To address this challenge, we propose Cognitive Graph-Field Memory (CGFM), a persistent multimodal scene representation that couples explicit relational memory with continuous spatial intuition. CGFM organizes objects, spatial relations, and visual observations into a multimodal scene graph, enabling target retrieval and long-horizon reasoning across navigation tasks. When no reliable target match is identified, graph-based evidence is projected into a goal-conditioned semantic-frontier field to guide exploration toward semantically promising frontiers and regions. Building upon CGFM, we introduce CGFM-Nav, a foundation-model-based framework for lifelong multimodal navigation that integrates task-relevant subgraph selection, VLM reasoning, and verification feedback into a closed decision loop. Preliminary experiments on GOAT-Bench show that, under the same Qwen3-VL-8B backbone, CGFM-Nav improves the overall success rate from 53.2% to 63.0% and SPL from 30.0% to 39.6%, demonstrating the effectiveness of combining explicit semantic memory with semantic-guided exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。