arXiv:2606.16935cs.ROcs.AI2026-06被引 1

让火星车用自然语言查询地图,实时识别可靠语义地标。

CrossMaps: Confidence-Aware Open-Vocabulary Semantic Mapping for Rover Navigation

论文配图:CrossMaps: Confidence-Aware Open-Vocabulary Semantic Mapping for Rover Navigation
图 1 · 摘自论文原文
  • 融合视觉、语义与时空置信度,动态筛选可靠观测
  • 构建短时/长时双记忆结构,持久化可信语义地标
  • 支持自然语言查询,适合移动机器人导航应用

火星车依赖感知构建包含物体和传感器质量信息(如测距可靠性、光照伪影、数据密度)的空间地图,以指导数据融合、嵌入更新及部分可观测下的导航。为研究这一耦合的感知-导航过程,我们提出CrossMaps,一种基于RGB-D数据的实时置信度感知开放词汇语义映射管道,可生成支持语言查询的地图。在VLMaps风格基础上,CrossMaps结合多尺度CLIP嵌入,采用置信度感知融合与双记忆架构(短时记忆STM与长时记忆LTM)。STM利用几何、语义与时间置信度聚合噪声视觉观测,将高置信且一致的单元提升至LTM作为持久语义地标。该系统部署于搭载Jetson Orin的无人地面载具(UGV)上,与SLAM协同运行,实现实时处理,并生成可被自然语言查询的语义热图,用于引导火星车导航。

原文摘要 · Abstract (English)

Rovers rely on perception to maintain spatial maps that encode both objects and sensor quality (e.g., range reliability, lighting artifacts, data density), guiding data fusion, embedding updates, and navigation under partial observability. To study these coupled perception-navigation processes, we present CrossMaps, a real-time confidence-aware open-vocabulary semantic mapping pipeline that constructs language-queryable maps from RGB-D data. Building on VLMaps-style approaches, CrossMaps integrates multi-scale CLIP embeddings with confidence-aware fusion and a dual-memory architecture consisting of Short-Term Memory (STM) and Long-Term Memory (LTM). The STM aggregates noisy visual observations using geometric, semantic, and temporal confidence cues, while confident and coherent cells are promoted to the LTM as persistent semantic landmarks. Designed for deployment with a Jetson Orin-powered UGV alongside SLAM, CrossMaps runs in real time and produces semantic heatmaps that can be queried with natural language to guide rover navigation.

语义地图开放词汇机器人导航置信度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。