arXiv:2506.23468cs.CV2025-06ICCV被引 49

让导航智能体像人一样持续学习环境变化,提升复杂场景下的语言导航能力。

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments

  • 用紧凑潜在表示建模环境动态,实现前瞻性规划与策略优化。
  • 在多个主流基准上显著提升导航成功率,表现优于现有方法。
  • 适合研究持续学习、视觉语言导航与具身智能的开发者参考。

在连续环境中的视觉-语言导航(VLN-CE)要求智能体根据自然语言指令执行一系列导航动作。当前方法常面临泛化到新环境及应对导航过程中环境变化的挑战。受人类认知启发,我们提出NavMorph——一种自演化世界模型框架,增强智能体对环境的理解与决策能力。该框架采用紧凑的潜在表示建模环境动态,赋予智能体前瞻性以实现自适应规划与策略改进。通过引入新型情境演化记忆(Contextual Evolution Memory),NavMorph利用场景上下文信息支持高效导航,同时保持在线适应性。大量实验表明,该方法在多个主流VLN-CE基准上取得显著性能提升。代码已公开于https://github.com/Feliciaxyao/NavMorph。

原文摘要 · Abstract (English)

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions. Current approaches often struggle with generalizing to novel environments and adapting to ongoing changes during navigation. Inspired by human cognition, we present NavMorph, a self-evolving world model framework that enhances environmental understanding and decision-making in VLN-CE tasks. NavMorph employs compact latent representations to model environmental dynamics, equipping agents with foresight for adaptive planning and policy refinement. By integrating a novel Contextual Evolution Memory, NavMorph leverages scene-contextual information to support effective navigation while maintaining online adaptability. Extensive experiments demonstrate that our method achieves notable performance improvements on popular VLN-CE benchmarks. Code is available at https://github.com/Feliciaxyao/NavMorph.

视觉语言导航自演化具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。