在线构建开放词汇全景地图,实时感知环境语义与几何。
OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian Splatting
- 用滑动窗口实现局部到全局的在线映射
- 通过几何语义融合,完整合并实例片段
- 支持开放词汇理解,适合机器人实时交互
开放词汇场景理解与在线全景映射对具身应用至关重要。然而现有方法多为离线或缺乏实例级理解,难以应用于真实机器人任务。本文提出OnlinePG,一种基于3D Gaussian Splatting的在线系统,融合几何重建与开放词汇感知。技术上,采用高效局部-全局范式与滑动窗口机制:构建3D语义聚类图,联合利用几何与语义线索,在滑动窗口内融合不一致段落形成完整实例;随后,将局部3D Gaussian地图以带空间属性的显式网格表示,通过鲁棒双向二分图匹配融合至全局地图;最后,利用空间网格中的融合视觉语言模型特征实现开放词汇场景理解。在多个主流数据集上的大量实验表明,该方法在在线方法中表现更优,且保持实时效率。
原文摘要 · Abstract (English)
Open-vocabulary scene understanding with online panoptic mapping is essential for embodied applications to perceive and interact with environments. However, existing methods are predominantly offline or lack instance-level understanding, limiting their applicability to real-world robotic tasks. In this paper, we propose OnlinePG, a novel and effective system that integrates geometric reconstruction and open-vocabulary perception using 3D Gaussian Splatting in an online setting. Technically, to achieve online panoptic mapping, we employ an efficient local-to-global paradigm with a sliding window. To build local consistency map, we construct a 3D segment clustering graph that jointly leverages geometric and semantic cues, fusing inconsistent segments within sliding window into complete instances. Subsequently, to update the global map, we construct explicit grids with spatial attributes for the local 3D Gaussian map and fuse them into the global map via robust bidirectional bipartite 3D Gaussian instance matching. Finally, we utilize the fused VLM features inside the 3D spatial attribute grids to achieve open-vocabulary scene understanding. Extensive experiments on widely used datasets demonstrate that our method achieves better performance among online approaches, while maintaining real-time efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。