让AI导航时能像人一样回忆过往经验,提升复杂环境下的导航能力。
CMMR-VLN: Vision-and-Language Navigation via Continual Multimodal Memory Retrieval
- 构建多模态记忆库,用视觉和地标索引过往导航经验
- 检索增强生成使模型可复用成功路径与错误教训,成功率最高提升200%
- 支持持续学习,适合长距离或陌生场景的智能导航任务
尽管大型语言模型(LLMs)被引入视觉-语言导航(VLN)以提升指令理解与泛化能力,但现有基于LLM的VLN缺乏选择性调用和利用先前经验的能力,限制了其在长时程和陌生场景中的表现。本文提出CMMR-VLN(基于持续多模态记忆检索的VLN),为LLM代理赋予结构化记忆与反思能力。具体而言,该框架构建以全景视觉图像和显著地标为索引的多模态经验记忆库,在导航过程中检索相关经验;引入检索增强生成流程,模拟资深人类导航者如何利用先验知识;并采用基于反思的记忆更新策略,仅存储完整成功路径及失败案例中的关键初始错误。综合实验表明,在仿真与真实测试中,相较于NavGPT、MapGPT和DiscussNav,平均成功率分别提升52.9%、20.9%、20.9%以及200%、50%、50%,凸显CMMR-VLN作为骨干VLN框架的巨大潜力。
原文摘要 · Abstract (English)
Although large language models (LLMs) are introduced into vision-and-language navigation (VLN) to improve instruction comprehension and generalization, existing LLM- based VLN lacks the ability to selectively recall and use relevant priori experiences to help navigation tasks, limiting their performance in long-horizon and unfamiliar scenarios. In this work, we propose CMMR-VLN (Continual Multimodal Memory Retrieval based VLN), a VLN framework that endows LLM agents with structured memory and reflection capabilities. Specifically, the CMMR-VLN constructs a multimodal experi- ence memory indexed by panoramic visual images and salient landmarks to retrieve relevant experiences during navigation, introduces a retrieved-augmented generation pipeline to mimick how experienced human navigators leverage priori knowledge, and incorporates a reflection-based memory update strategy that selectively stores complete successful paths and the key initial mistake in failure cases. Comprehensive tests illustrate average success rate improvements of 52.9%, 20.9% and 20.9%, and 200%, 50% and 50% over the NavGPT, the MapGPT, and the DiscussNav in simulation and real tests, respectively eluci- dating the great potential of the CMMR-VLN as a backbone VLN framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。