用手绘地图引导智能体导航未知环境,突破传统路径依赖。
SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps
- 基于手绘地图与视觉观测对齐,实现抽象地图下的精准导航
- 在高抽象度地图上相对提升SPL 105%,显著优于现有方法
- 适合关注具身智能、地图理解与人机交互的研究者
人类常通过手绘路线图指导导航。受此启发,我们提出基于手绘地图的具身导航任务SkeNa,要求智能体仅凭手绘草图在未见过的环境中抵达目标。为此构建了大规模数据集SoR,包含71个室内场景中54,000组轨迹与手绘地图对。SoR设立两个不同抽象层级的验证集,依据草图对空间尺度的保留程度分类。我们开发自动化草图生成管道,将平面图高效转换为手绘风格。针对SkeNa,提出SkeNavigator框架:利用射线式地图描述符(RMD)增强草图特征表示;通过双地图对齐的目标预测器(DAGP),结合现场探索地图与草图特征对应关系预测目标位置并引导导航。SkeNavigator在高抽象验证集上相对提升SPL达105%。代码与数据集将公开。
原文摘要 · Abstract (English)
A typical human strategy for giving navigation guidance is to sketch route maps based on the environmental layout. Inspired by this, we introduce Sketch map-based visual Navigation (SkeNa), an embodied navigation task in which an agent must reach a goal in an unseen environment using only a hand-drawn sketch map as guidance. To support research for SkeNa, we present a large-scale dataset named SoR, comprising 54k trajectory and sketch map pairs across 71 indoor scenes. In SoR, we introduce two navigation validation sets with varying levels of abstraction in hand-drawn sketches, categorized based on their preservation of spatial scales in the environment, to facilitate future research. To construct SoR, we develop an automated sketch-generation pipeline that efficiently converts floor plans into hand-drawn representations. To solve SkeNa, we propose SkeNavigator, a navigation framework that aligns visual observations with hand-drawn maps to estimate navigation targets. It employs a Ray-based Map Descriptor (RMD) to enhance sketch map valid feature representation using equidistant sampling points and boundary distances. To improve alignment with visual observations, a Dual-Map Aligned Goal Predictor (DAGP) leverages the correspondence between sketch map features and on-site constructed exploration map features to predict goal position and guide navigation. SkeNavigator outperforms prior floor plan navigation methods by a large margin, improving SPL on the high-abstract validation set by 105% relatively. Our code and dataset will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。