arXiv:2605.17249cs.RO2026-05

提出双系统框架,提升视觉语言导航的空间感知与规划效率。

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation

论文配图:SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation
图 1 · 摘自论文原文
  • 双系统设计:一个具空间感知的VLM生成动作,另一个用地图与图像规划路径。
  • 在VLN-CE上达到最新最好效果,长程导航更稳定。
  • 适合需要高精度导航与快速决策的应用场景。

视觉语言导航(VLN)目前主要采用两种范式:端到端的视觉语言模型(VLM)通过导航轨迹微调直接预测动作;以及无需训练的零样本模块化流程,利用预训练多模态大语言模型(MLLM)实现对未见环境的泛化。然而,端到端方法在长程导航中表现不佳且缺乏动态推理能力,而零样本方法受限于空间定位不足,且推理耗时较长。为此,我们提出SEDualVLN,一种增强空间感知的双系统VLN框架。系统1是融合全局与局部空间意识的VLM,用于动作生成;系统2结合通用MLLM与地图模块,通过实时3D地图的俯视图和渲染路径图像,由MLLM规划航点。两个系统分别通过不同形式的空间增强提升导航方向感,最终协同完成任务。该框架采用快慢配合策略,显著提升导航性能。在VLN-CE基准测试中达到当前最优表现,消融实验验证了各系统与模块的有效性。

原文摘要 · Abstract (English)

Vision-Language Navigation (VLN) approaches have currently followed two primary paradigms: the end-to-end Vision-Language Model (VLM) policy fine-tuned on navigation trajectories to directly predict actions, and the zero-shot modular pipeline integrating pre-trained Multimodal Large Language Model (MLLM) for training-free generalization to unseen environments. However, end-to-end methods struggle with long-horizon navigation and lack dynamic reasoning, whereas zero-shot methods are constrained by limited spatial grounding for reliable planning and also require substantial reasoning time. To bridge this gap, we introduce SEDualVLN, a spatially-enhanced dual-system VLN framework. System 1 is a VLM model enhanced with both global and local spatial awareness, used for action generation. System 2 integrates a general MLLM with a mapping module, wherein the MLLM plans waypoints by leveraging top-down views of the real-time 3D map alongside streams of rendered path images. Both systems leverage different forms of spatial enhancement to cultivate the agent's sense of direction in VLN tasks. Ultimately, they cooperate to complete the navigation task through a fast-slow coordinated approach. SEDualVLN achieves state-of-the-art performance on VLN-CE benchmarks, and further ablation studies demonstrate the effectiveness of each system and module.

视觉导航多模态路径规划大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。