提出StereoNav框架,提升视觉语言导航在真实场景的稳定性。
What Limits Vision-and-Language Navigation ?

- 用目标位置先验增强跨域空间定位能力
- 通过双目视觉融合语义与几何信息,提升深度感知
- 少参数少数据下实测性能超越主流方法
视觉-语言导航(VLN)是具身智能的核心。当前代理在从仿真到真实部署时性能显著下降,主要源于感知不稳定性(如光照变化、运动模糊)和指令不明确。现有方法通过扩大模型规模和训练数据来缩小差距,但我们认为瓶颈在于缺乏鲁棒的空间定位和跨域先验。本文提出StereoNav,一种增强真实世界导航一致性的视觉-语言-动作框架。为弥合合成训练与物理执行的差距,引入目标位置先验作为持久桥梁,提供跨域不变的视觉引导,有效在指令模糊时实现定位。同时,利用双目视觉构建语义与几何统一表征,通过增强深度感知实现精准动作预测,缓解运动模糊和光照漂移。在R2R-CE和RxR-CE上的实验表明,StereoNav达到最先进的自视点RGB性能,SR和SPL分别为81.1%和68.3%,以及67.5%和52.0%,且参数量和训练数据远低于以往扩增方法。更重要的是,真实机器人部署验证了其在复杂非结构化环境中的导航可靠性。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) is a cornerstone of embodied intelligence. However, current agents often suffer from significant performance degradation when transitioning from simulation to real-world deployment, primarily due to perceptual instability (e.g., lighting variations and motion blur) and under-specified instructions. While existing methods attempt to bridge this gap by scaling up model size and training data, we argue that the bottleneck lies in the lack of robust spatial grounding and cross-domain priors. In this paper, we propose StereoNav, a robust Vision-Language-Action framework designed to enhance real-world navigation consistency. To address the inherent gap between synthetic training and physical execution, we introduce Target-Location Priors as a persistent bridge. These priors provide stable visual guidance that remains invariant across domains, effectively grounding the agent even when instructions are vague. Furthermore, to mitigate visual disturbances like motion blur and illumination shifts, StereoNav leverages stereo vision to construct a unified representation of semantics and geometry, enabling precise action prediction through enhanced depth awareness. Extensive experiments on R2R-CE and RxR-CE demonstrate that StereoNav achieves state-of-the-art egocentric RGB performance, with SR and SPL scores of 81.1% and 68.3%, and 67.5% and 52.0%, respectively, while using significantly fewer parameters and less training data than prior scaling-based approaches. More importantly, real-world robotic deployments confirm that StereoNav substantially improves navigation reliability in complex, unstructured environments. Project page: https://yunheng-wang.github.io/stereonav-public.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。