arXiv:2606.14267cs.RO2026-06中稿 · CVPR

用楼图引导多模态导航,提升机器人在新环境中的探索效率。

FloVerse: Floor Plan-Guided Multi-Modal Navigation

论文配图:FloVerse: Floor Plan-Guided Multi-Modal Navigation
图 1 · 摘自论文原文
  • 构建统一任务框架,融合目标点、物体和图像三种导航方式。
  • 在1.6万场景数据集上实现24万条专家轨迹与1200万帧视觉数据。
  • 提出ThreeDiff模型,通过扩散机制隐式利用楼图空间信息。

楼图蕴含紧凑的空间先验,使智能体能更高效地导航未见过的场景。现有研究主要聚焦于点导航(PointNav)及有限环境。为弥补这一差距,我们提出FloVerse,一项统一点导航、物体导航和图像导航的新任务。为此,我们构建了包含1.6K个场景的大型数据集FloVerse-1.6K,源自HM3D与Gibson 4+,配以对应楼图,包含240万条专家轨迹和1200万帧RGBD图像。我们进一步提出ThreeDiff,一种两阶段模仿学习策略,包含规划器、基于扩散的多模态目标推理模块(通过掩码模态建模训练),以及基于深度的轨迹精炼模块,用于安全执行。大量实验表明:(1) 楼图先验在所有目标模态下均提升导航性能;(2) ThreeDiff隐式捕捉楼图中的空间信息。这些结果验证了空间先验的有效性,并支持我们提出的统一楼图引导具身导航方法。

原文摘要 · Abstract (English)

Floor plans encapsulate compact spatial priors, enabling agents to navigate unseen scenes more efficiently. While prior work has explored floor plan-guided navigation, it has focused mainly on PointNav and a limited set of environments. To bridge this gap, we introduce FloVerse, a new task for floor plan-guided embodied navigation that unifies PointNav, ObjectNav, and ImageNav. To support FloVerse, we assemble FloVerse-1.6K, a large-scale dataset of 1.6K scenes from HM3D and Gibson 4+, paired with corresponding floor plans, comprising 240K expert trajectories and 12M RGBD frames. We further propose ThreeDiff, a two-stage imitation learning policy comprising a planner, a diffusion-based multimodal goal-reasoning module trained via masked-modality modeling, and a refiner, a depth-based trajectory-refinement module for safe execution. Extensive experiments demonstrate that (1) floor-plan priors improve navigation performance across all goal modalities, and (2) ThreeDiff implicitly captures spatial information from floor plans. These results underscore the effectiveness of spatial priors and validate our proposed unified approach for floor plan-guided embodied navigation.

具身导航楼图引导多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。