对比虚拟与现实环境下的视觉语言导航表现,揭示现有方法的落地差距。
A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation

- 按动作与模型框架分类梳理主流导航方法
- 实测显示真实场景成功率比仿真低近40个百分点
- 建议关注感知、决策与控制的协同优化
导航是自主系统的基本能力,但现有方法多依赖高度结构化模型和强先验假设,难以在开放、不确定的真实环境中保持鲁棒性。视觉-语言导航(VLN)通过数据驱动方式融合自然语言理解与视觉感知,展现出广阔前景。尽管研究关注度不断提升,但系统性的方法分类与真实世界验证仍较缺乏。本综述对VLN研究进行了全面回顾,从两个正交维度组织最新方法:动作范式(分层与单体框架)和模型范式(判别式与生成式)。分析了各类方法的优劣。此外,在物理机器人平台上对代表性系统配置进行了系统性真实世界评估。在十种不同真实场景中测试发现,模拟与真实部署间存在显著性能差距:一种典型的单体RGB仅方法在仿真中成功率达61%,但在真实环境中降至22%;而分层框架在真实环境中实现51%的成功率,表现出更强鲁棒性。最后,指出了感知、决策与控制方面的关键挑战,需在未来研究中重点突破。
原文摘要 · Abstract (English)
Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness in open and uncertain real-world environments. Vision-and-Language Navigation (VLN) offers a promising direction by enabling robots to integrate natural language understanding with visual perception in a data-driven manner. Although VLN has attracted increasing research attention, systematic methodological taxonomy and real-world validation remain limited. This survey presents a comprehensive review of VLN research. Specifically, state-of-the-art methods are organized along two orthogonal dimensions: action paradigms, including hierarchical and monolithic frameworks, and model paradigms, including discriminative and generative approaches. A critical analysis of their respective strengths and limitations is provided. Additionally, we conduct a systematic real-world evaluation of representative VLN system configurations on a physical robotic platform. Experiments across ten diverse real-world scenes show a substantial performance gap between simulation and real-world deployment under the tested configurations: a representative monolithic RGB-only method achieves 61% success in simulation but drops to 22% in real-world deployment, while a hierarchical framework achieves a higher real-world success rate of 51%, suggesting stronger robustness in our evaluation setting. Finally, we highlight key challenges in perception, decision-making, and control that must be addressed in future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。