用特征场实现稀疏语义信号下的高精度3D重建与实时渲染
GSFF-SLAM: 3D Semantic Gaussian Splatting SLAM via Feature Field
- 基于3D高斯点云和特征场联合优化,支持多模态2D语义先验
- 在真实场景下达到95.03% mIoU,速度提升2.9倍且性能损失小
- 适合需要高效语义重建的自动驾驶与机器人导航系统
语义感知的3D场景重建对自主机器人完成复杂交互至关重要。语义SLAM作为一种在线方法,将位姿追踪、几何重建与语义映射统一建模,潜力巨大。然而,现有系统依赖2D真值先验进行监督,常受限于真实环境中信号稀疏和噪声问题。为此,我们提出GSFF-SLAM,一种基于3D高斯点云的新型稠密语义SLAM系统,利用特征场实现外观、几何与N维语义特征的联合渲染。通过独立优化特征梯度,本方法支持多种形式的2D先验,尤其适用于稀疏且嘈杂的信号。实验表明,该方法在追踪精度与照片级渲染质量上均优于现有方法。使用2D真值先验时,达到95.03% mIoU的语义分割性能,同时实现最高2.9倍的速度提升,仅伴随轻微性能下降。
原文摘要 · Abstract (English)
Semantic-aware 3D scene reconstruction is essential for autonomous robots to perform complex interactions. Semantic SLAM, an online approach, integrates pose tracking, geometric reconstruction, and semantic mapping into a unified framework, shows significant potential. However, existing systems, which rely on 2D ground truth priors for supervision, are often limited by the sparsity and noise of these signals in real-world environments. To address this challenge, we propose GSFF-SLAM, a novel dense semantic SLAM system based on 3D Gaussian Splatting that leverages feature fields to achieve joint rendering of appearance, geometry, and N-dimensional semantic features. By independently optimizing feature gradients, our method supports semantic reconstruction using various forms of 2D priors, particularly sparse and noisy signals. Experimental results demonstrate that our approach outperforms previous methods in both tracking accuracy and photorealistic rendering quality. When utilizing 2D ground truth priors, GSFF-SLAM achieves state-of-the-art semantic segmentation performance with 95.03\% mIoU, while achieving up to 2.9$\times$ speedup with only marginal performance degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。