用楼层图引导导航,让智能体用简短指令完成复杂路径规划。
FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation
- 引入楼层图作为全局空间先验,结合简洁指令进行导航。
- 在超1万条路径上测试,导航成功率提升60%以上。
- 适合需要空间理解能力的智能体导航研究者使用。
现有视觉语言导航任务依赖冗长指令,忽略有用的全局空间信息,限制了对空间结构的推理能力。尽管真实建筑中普遍存在可读的平面图,但现有智能体缺乏理解与利用此类信息的能力。为此,我们提出新范式 FloorPlan-VLN,将结构化语义楼层图作为全局空间先验,实现仅凭简洁指令的导航。我们构建了包含72个场景、超过10,000个智能体轨迹的 FloorPlan-VLN 数据集,配对超过100张语义标注的楼层图与 Matterport3D 基础的导航路径及省略步骤指引的简短指令。同时提出简单有效的 FP-Nav 方法,采用双视角时空对齐视频序列与辅助推理任务,实现观测、楼层图与指令的一致对齐。在新基准下评估,该方法显著优于适配的先进 VLN 基线,导航成功率达到相对提升60%以上。此外,通过噪声建模与真实部署验证,证明了 FP-Nav 对执行漂移和楼层图失真的鲁棒性。结果验证了基于楼层图导航的有效性,并表明 FloorPlan-VLN 是迈向更智能空间导航的重要一步。
原文摘要 · Abstract (English)
Existing Vision-Language Navigation (VLN) task requires agents to follow verbose instructions, ignoring some potentially useful global spatial priors, limiting their capability to reason about spatial structures. Although human-readable spatial schematics (e.g., floor plans) are ubiquitous in real-world buildings, current agents lack the cognitive ability to comprehend and utilize them. To bridge this gap, we introduce \textbf{FloorPlan-VLN}, a new paradigm that leverages structured semantic floor plans as global spatial priors to enable navigation with only concise instructions. We first construct the FloorPlan-VLN dataset, which comprises over 10k episodes across 72 scenes. It pairs more than 100 semantically annotated floor plans with Matterport3D-based navigation trajectories and concise instructions that omit step-by-step guidance. Then, we propose a simple yet effective method \textbf{FP-Nav} that uses a dual-view, spatio-temporally aligned video sequence, and auxiliary reasoning tasks to align observations, floor plans, and instructions. When evaluated under this new benchmark, our method significantly outperforms adapted state-of-the-art VLN baselines, achieving more than a 60\% relative improvement in navigation success rate. Furthermore, comprehensive noise modeling and real-world deployments demonstrate the feasibility and robustness of FP-Nav to actuation drift and floor plan distortions. These results validate the effectiveness of floor plan guided navigation and highlight FloorPlan-VLN as a promising step toward more spatially intelligent navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。