让机器人逐步理解3D场景,支持动态更新语义与关系图谱。
OGScene3D: Incremental Open-Vocabulary 3D Gaussian Scene Graph Mapping for Scene Understanding
- 用置信度加权的高斯表示法,同时建模语义和可靠性。
- 通过局部对齐与全局优化,实现跨时间的语义一致性。
- 适合需要持续感知新环境的机器人应用,如导航与操作。
开放词汇场景理解对机器人应用至关重要,使机器人能够理解复杂的三维环境上下文,并支持导航、操作等下游任务。然而,现有方法需预先构建完整的3D语义地图才能生成场景图,难以适应机器人在探索中逐步构建环境的现实场景。为此,我们提出OGScene3D,一种可增量进行3D语义映射与场景图构建的开放词汇理解系统。该系统采用基于置信度的高斯语义表示,联合建模语义预测及其可靠性,提升场景建模鲁棒性。在此基础上,设计分层3D语义优化策略,通过局部对应建立与全局精炼实现语义一致性,构建全局一致的语义地图。此外,引入长期全局优化方法,利用历史观测的时间记忆增强语义预测。结合2D-3D语义一致性与高斯渲染贡献,持续优化整个场景的语义理解。进一步开发渐进式图谱构建方法,动态创建与更新节点及语义关系,支持3D场景图的连续更新。在广泛使用的数据集与真实场景上的大量实验验证了OGScene3D在开放词汇场景理解中的有效性。
原文摘要 · Abstract (English)
Open-vocabulary scene understanding is crucial for robotic applications, enabling robots to comprehend complex 3D environmental contexts and supporting various downstream tasks such as navigation and manipulation. However, existing methods require pre-built complete 3D semantic maps to construct scene graphs for scene understanding, which limits their applicability in robotic scenarios where environments are explored incrementally. To address this challenge, we propose OGScene3D, an open-vocabulary scene understanding system that achieves accurate 3D semantic mapping and scene graph construction incrementally. Our system employs a confidence-based Gaussian semantic representation that jointly models semantic predictions and their reliability, enabling robust scene modeling. Building on this representation, we introduce a hierarchical 3D semantic optimization strategy that achieves semantic consistency through local correspondence establishment and global refinement, thereby constructing globally consistent semantic maps. Moreover, we design a long-term global optimization method that leverages temporal memory of historical observations to enhance semantic predictions. By integrating 2D-3D semantic consistency with Gaussian rendering contribution, this method continuously refines the semantic understanding of the entire scene. Furthermore, we develop a progressive graph construction approach that dynamically creates and updates both nodes and semantic relationships, allowing continuous updating of the 3D scene graphs. Extensive experiments on widely used datasets and real-world scenes demonstrate the effectiveness of our OGScene3D on open-vocabulary scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。