用DINO+几何编码提升3D SLAM的语义与几何理解能力
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
- 融合DINO语义与几何编码器生成带几何感知的特征
- 在Replica、ScanNet、TUM数据集上优于现有方法
- 适合做神经隐式/显式3D重建的开发者参考
本文提出DINO-SLAM,一种基于DINO的增强策略,通过引入更全面的语义理解来提升SLAM系统中神经隐式(NeRF)与显式表示(高斯点云渲染——GS)的性能。然而,仅依赖DINO本身缺乏对3D几何结构的理解,仅带来有限改进。为此,我们设计了场景几何编码器(SGE),将原始DINO特征转化为具有几何感知的geoDINO特征,以捕捉原生DINO无法建模的几何关系。在此基础上,我们构建了两种集成geoDINO特征的NeRF与GS SLAM基础范式。在Replica、ScanNet和TUM数据集上的实验表明,相较于最先进方法,我们的DINO增强管道取得了更优表现。
原文摘要 · Abstract (English)
This paper presents DINO-SLAM, a DINO-informed design strategy to enhance implicit (Neural Radiance Field -- NeRF) and explicit representations (Gaussian Splatting -- GS) in SLAM systems through the more comprehensive semantics understanding enabled by DINO. This latter alone, however, lacks proper 3D geometry understanding, allowing only for marginal improvements. Therefore, we rely on a Scene Geometry Encoder (SGE) to enrich DINO features into geometry-aware DINO features (geoDINO), to better understand those geometric relationships that vanilla DINO features fail to capture. Building upon it, we propose two foundational paradigms for NeRF and GS SLAM systems integrating geoDINO features. Compared to state-of-the-art methods, our DINO-informed pipelines achieve superior performance on the Replica, ScanNet, and TUM datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。