让视觉语言模型在三维空间中实现可量化不确定性的开放词典映射。
LatentBKI: Open-Dictionary Continuous Mapping in Visual-Language Latent Spaces with Quantifiable Uncertainty
- 基于贝叶斯核推断,将视觉语言嵌入动态融合进体素地图
- 在Matterport3D和Semantic KITTI上实现开放词汇查询,不确定性可量化
- 适合需要灵活语义理解的机器人导航场景
本文提出一种新型概率映射算法LatentBKI,支持可量化不确定性的开放词汇映射。传统语义映射方法局限于固定语义类别,限制了其在复杂机器人任务中的应用。视觉-语言(VL)模型通过在潜在空间联合建模语言与视觉特征,突破了预定义类别的限制。LatentBKI将VL模型的神经嵌入递归融入体素地图,利用邻近观测的空间相关性,通过贝叶斯核推断(BKI)实现不确定性量化。在Matterport3D和Semantic KITTI数据集上的实验表明,LatentBKI在保持连续映射概率优势的同时,支持开放词典查询。真实世界实验验证了其在复杂室内环境中的适用性。
原文摘要 · Abstract (English)
This paper introduces a novel probabilistic mapping algorithm, LatentBKI, which enables open-vocabulary mapping with quantifiable uncertainty. Traditionally, semantic mapping algorithms focus on a fixed set of semantic categories which limits their applicability for complex robotic tasks. Vision-Language (VL) models have recently emerged as a technique to jointly model language and visual features in a latent space, enabling semantic recognition beyond a predefined, fixed set of semantic classes. LatentBKI recurrently incorporates neural embeddings from VL models into a voxel map with quantifiable uncertainty, leveraging the spatial correlations of nearby observations through Bayesian Kernel Inference (BKI). LatentBKI is evaluated against similar explicit semantic mapping and VL mapping frameworks on the popular Matterport3D and Semantic KITTI datasets, demonstrating that LatentBKI maintains the probabilistic benefits of continuous mapping with the additional benefit of open-dictionary queries. Real-world experiments demonstrate applicability to challenging indoor environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。