用平面图指导视觉导航,提升陌生环境下的定位效率与准确性。
FloNa: Floor Plan Guided Embodied Visual Navigation
- 引入平面图引导的视觉导航框架FloNa,结合扩散模型与定位模块对齐图像与平面图。
- 在iGibson中构建2万条导航轨迹,验证了在陌生场景下使用平面图可显著提升导航效果。
- 适合做智能机器人、自动驾驶等需要地图辅助的视觉导航研究者参考。
人类在不熟悉环境中常依赖平面图进行导航,因其易获取、可靠且提供丰富的几何信息。然而,现有视觉导航方法忽视这一重要先验知识,导致效率和精度受限。为此,我们提出首个融合平面图的具身视觉导航任务——FloNa。该任务面临两大挑战:一是处理平面图与实际场景布局之间的空间不一致以实现无碰撞导航;二是对齐观测图像与平面图草图,克服二者模态差异。为此,我们提出FloDiff,一种结合定位模块的扩散策略框架,有效实现当前观测与平面图的对齐。我们还在iGibson模拟器中收集了20,000条导航轨迹,覆盖117个场景,用于训练与评估。大量实验表明,本框架在利用平面图知识时,在陌生场景中展现出显著更高的有效性与效率。
原文摘要 · Abstract (English)
Humans naturally rely on floor plans to navigate in unfamiliar environments, as they are readily available, reliable, and provide rich geometrical guidance. However, existing visual navigation settings overlook this valuable prior knowledge, leading to limited efficiency and accuracy. To eliminate this gap, we introduce a novel navigation task: Floor Plan Visual Navigation (FloNa), the first attempt to incorporate floor plan into embodied visual navigation. While the floor plan offers significant advantages, two key challenges emerge: (1) handling the spatial inconsistency between the floor plan and the actual scene layout for collision-free navigation, and (2) aligning observed images with the floor plan sketch despite their distinct modalities. To address these challenges, we propose FloDiff, a novel diffusion policy framework incorporating a localization module to facilitate alignment between the current observation and the floor plan. We further collect $20k$ navigation episodes across $117$ scenes in the iGibson simulator to support the training and evaluation. Extensive experiments demonstrate the effectiveness and efficiency of our framework in unfamiliar scenes using floor plan knowledge. Project website: https://gauleejx.github.io/flona/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。