提出新导航框架,让AI在无训练数据下也能听懂指令走迷宫。
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
- 将导航拆解为带约束的子指令,逐步完成任务
- 在两个基准上成功率分别提升12%和13%
- 适用于真实机器人部署,跨场景表现强
我们解决连续环境中的零样本视觉语言导航(VLN-CE)问题。由于缺乏专家示范且环境结构先验信息极少,该任务极具挑战性。为此,我们提出约束感知导航器(CA-Nav),将零样本VLN-CE重构为序列化的、约束感知的子指令补全过程。CA-Nav通过两个核心模块持续将子指令转化为导航计划:约束感知子指令管理器(CSM)定义分解后子指令的完成条件作为约束,并以约束驱动方式切换子指令以跟踪导航进度;约束感知价值映射器(CVM)在CSM约束指导下实时生成价值图,并利用超像素聚类进行优化以提升导航稳定性。CA-Nav在两个VLN-CE基准上达到当前最优性能,在R2R-CE和RxR-CE的验证未见分割上,成功率分别较之前最佳方法提升12%和13%。此外,该方法在多种室内场景的真实机器人部署中也表现出显著有效性。
原文摘要 · Abstract (English)
We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting. Zero-shot VLN-CE is particularly challenging due to the absence of expert demonstrations for training and minimal environment structural prior to guide navigation. To confront these challenges, we propose a Constraint-Aware Navigator (CA-Nav), which reframes zero-shot VLN-CE as a sequential, constraint-aware sub-instruction completion process. CA-Nav continuously translates sub-instructions into navigation plans using two core modules: the Constraint-Aware Sub-instruction Manager (CSM) and the Constraint-Aware Value Mapper (CVM). CSM defines the completion criteria for decomposed sub-instructions as constraints and tracks navigation progress by switching sub-instructions in a constraint-aware manner. CVM, guided by CSM's constraints, generates a value map on the fly and refines it using superpixel clustering to improve navigation stability. CA-Nav achieves the state-of-the-art performance on two VLN-CE benchmarks, surpassing the previous best method by 12 percent and 13 percent in Success Rate on the validation unseen splits of R2R-CE and RxR-CE, respectively. Moreover, CA-Nav demonstrates its effectiveness in real-world robot deployments across various indoor scenes and instructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。