一次性构建兼具视觉、几何与语义的3D地图,实时高效且精度高。
OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics
- 融合3DGS与体素的混合表示,兼顾细节与结构稳定。
- 在多个场景中实现渲染保真度、几何精度与零样本分割领先表现。
- 适合需要全面环境理解的机器人导航与交互任务。
机器人系统需要精确且全面的3D环境感知,要求同时捕捉照片级真实感外观(光学)、精确布局形状(几何)和开放词汇场景理解(语义)。现有方法通常仅部分满足这些需求,且存在光学模糊、几何不规则和语义歧义问题。为此,我们提出OmniMap。OmniMap是首个在线映射框架,能同步捕捉光学、几何和语义场景属性,同时保持实时性能和模型紧凑性。在架构层面,采用紧密耦合的3DGS-Voxel混合表示,结合精细建模与结构稳定性。在实现层面,针对不同模态的关键挑战引入多项创新:自适应相机建模以补偿运动模糊与曝光差异,结合法向量约束的混合增量表示,以及概率融合以实现鲁棒的实例级理解。大量实验表明,相比当前最优方法,OmniMap在多种场景下均展现出更优的渲染保真度、几何精度和零样本语义分割能力。其多功能性还通过多领域场景问答、交互式编辑、感知引导操作和地图辅助导航等下游应用得到验证。
原文摘要 · Abstract (English)
Robotic systems demand accurate and comprehensive 3D environment perception, requiring simultaneous capture of photo-realistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic). Existing methods typically achieve only partial fulfillment of these requirements while exhibiting optical blurring, geometric irregularities, and semantic ambiguities. To address these challenges, we propose OmniMap. Overall, OmniMap represents the first online mapping framework that simultaneously captures optical, geometric, and semantic scene attributes while maintaining real-time performance and model compactness. At the architectural level, OmniMap employs a tightly coupled 3DGS-Voxel hybrid representation that combines fine-grained modeling with structural stability. At the implementation level, OmniMap identifies key challenges across different modalities and introduces several innovations: adaptive camera modeling for motion blur and exposure compensation, hybrid incremental representation with normal constraints, and probabilistic fusion for robust instance-level understanding. Extensive experiments show OmniMap's superior performance in rendering fidelity, geometric accuracy, and zero-shot semantic segmentation compared to state-of-the-art methods across diverse scenes. The framework's versatility is further evidenced through a variety of downstream applications, including multi-domain scene Q&A, interactive editing, perception-guided manipulation, and map-assisted navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。