arXiv:2507.13152cs.CVcs.AI2025-07被引 13

让导航模型像生物一样不断进化,靠经验自我提升。

SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models

  • 构建三级模块:记忆、检索推理、反思,实现边走边学
  • 在未知环境成功率达57%和35.2%,比现有方法高15%-24%
  • 适合研究持续学习与智能体自演化的人看

视觉语言导航(VLN)近年发展得益于大语言模型(LLM),其在指令理解与任务推理上表现优异。但受限于固定知识库与推理能力,难以融入实际经验,缺乏高效演化能力。为此,受自然智能体启发,我们提出首个基于多模态大模型的自演化视觉语言导航框架(SE-VLN)。该框架包含三个核心模块:分层记忆模块将成功与失败案例转化为可复用知识;检索增强型思维推理模块通过调用经验支持多步决策;反思模块实现持续演化。实验证明,SE-VLN在未见过的R2R与REVERSE数据集上分别达到57%和35.2%的导航成功率,相比当前最优方法绝对提升23.9%和15.0%。且随着经验库增大,性能持续提升,展现出强大自演化潜力。

原文摘要 · Abstract (English)

Recent advances in vision-language navigation (VLN) were mainly attributed to emerging large language models (LLMs). These methods exhibited excellent generalization capabilities in instruction understanding and task reasoning. However, they were constrained by the fixed knowledge bases and reasoning abilities of LLMs, preventing fully incorporating experiential knowledge and thus resulting in a lack of efficient evolutionary capacity. To address this, we drew inspiration from the evolution capabilities of natural agents, and proposed a self-evolving VLN framework (SE-VLN) to endow VLN agents with the ability to continuously evolve during testing. To the best of our knowledge, it was the first time that an multimodal LLM-powered self-evolving VLN framework was proposed. Specifically, SE-VLN comprised three core modules, i.e., a hierarchical memory module to transfer successful and failure cases into reusable knowledge, a retrieval-augmented thought-based reasoning module to retrieve experience and enable multi-step decision-making, and a reflection module to realize continual evolution. Comprehensive tests illustrated that the SE-VLN achieved navigation success rates of 57% and 35.2% in unseen environments, representing absolute performance improvements of 23.9% and 15.0% over current state-of-the-art methods on R2R and REVERSE datasets, respectively. Moreover, the SE-VLN showed performance improvement with increasing experience repository, elucidating its great potential as a self-evolving agent framework for VLN.

视觉导航自演化多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。