arXiv:2605.03669cs.ROcs.AI2026-05

通过融合体素与实例层,实现可扩展的高精度开放词汇语义地图构建。

FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping by 3D Fusion of Voxel- and Instance-Level Layers

论文配图:FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping by 3D Fusion of Voxel- and Instance-Level Layers
图 1 · 摘自论文原文
  • 双层结构同步维护体素级与实例级语义层,共享同一体素地图。
  • 跨层语义融合提升两层质量,支持多层楼规模的精准映射。
  • 无需训练,适用于大规模场景,适合机器人空间认知任务。

开放词汇语义映射使机器人能在不依赖预定义类别的情况下,对未见过的概念进行空间定位。现有无训练方法通常通过多视角语义嵌入融合构建3D地图,或在实例级别分割图像并编码片段,或直接将图像块嵌入投影至稠密语义地图。后者避免了分割和2D到3D实例关联,但现有方法在可扩展性上受限。本文提出FUS3DMaps,一种在线双层语义映射方法,在共享体素地图中同时维护密集层与实例级层。该设计支持两层嵌入在体素级别进一步融合,结合两种方法的优势。实验表明,所提出的跨层融合策略提升了实例级与密集层的质量,且通过空间滑动窗口限制,实现了高效可扩展的实例级地图。在主流3D语义分割基准及多个大规模场景上的测试显示,FUS3DMaps可在多层建筑尺度下实现高精度开放词汇语义映射。补充材料与代码将公开:https://githanonymous.github.io/FUS3DMaps/

原文摘要 · Abstract (English)

Open-vocabulary semantic mapping enables robots to spatially ground previously unseen concepts without requiring predefined class sets. Current training-free methods commonly rely on multi-view fusion of semantic embeddings into a 3D map, either at the instance-level via segmenting views and encoding image crops of segments, or by projecting image patch embeddings directly into a dense semantic map. The latter approach sidesteps segmentation and 2D-to-3D instance association by operating on full uncropped image frames, but existing methods remain limited in scalability. We present FUS3DMaps, an online dual-layer semantic mapping method that jointly maintains both dense and instance-level open-vocabulary layers within a shared voxel map. This design enables further voxel-level semantic fusion of the layer embeddings, combining the complementary strengths of both semantic mapping approaches. We find that our proposed semantic cross-layer fusion approach improves the quality of both the instance-level and dense layers, while also enabling a scalable and highly accurate instance-level map where the dense layer and cross-layer fusion are restricted to a spatial sliding window. Experiments on established 3D semantic segmentation benchmarks as well as a selection of large-scale scenes show that FUS3DMaps achieves accurate open-vocabulary semantic mapping at multi-story building scales. Additional material and code will be made available: https://githanonymous.github.io/FUS3DMaps/.

语义地图开放词汇3D融合机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。