arXiv:2501.07015cs.CV2025-01被引 10

用SLAM引导的高斯点云实现单目实时高清三维重建

SplatMAP: Online Dense Monocular SLAM with 3D Gaussian Splatting

  • 结合SLAM稠密点云动态优化高斯模型密度
  • 融合几何约束与光照一致性,提升重建精度
  • 适合需要实时高清建模的机器人导航场景

从单目视频实现高保真三维重建仍具挑战,传统SfM和单目SLAM难以精确捕捉细节。尽管可微渲染如NeRF有所改善,但计算开销大,不适用于实时应用。现有3D高斯喷溅(3DGS)方法多关注光照一致性,忽略几何准确性,且未利用SLAM的动态深度与位姿更新来优化场景。本文提出一种融合稠密SLAM与3DGS的框架,实现实时高保真重建。引入SLAM感知自适应增密机制,利用SLAM生成的稠密点云动态更新并增加高斯模型密度;同时采用几何引导优化,结合边缘感知几何约束与光照一致性,联合优化3DGS的外观与几何结构,实现精细准确的映射。在Replica和TUM-RGBD数据集上的实验表明,本方法优于现有单目系统:在Replica上PSNR达36.864,SSIM为0.985,LPIPS为0.040,分别提升10.7%、6.4%和49.4%;在TUM-RGBD上相较最接近基线提升10.2%、6.6%和34.7%。结果验证了该框架在弥合光照与几何表示差距方面的潜力,推动实用高效的单目稠密重建发展。

原文摘要 · Abstract (English)

Achieving high-fidelity 3D reconstruction from monocular video remains challenging due to the inherent limitations of traditional methods like Structure-from-Motion (SfM) and monocular SLAM in accurately capturing scene details. While differentiable rendering techniques such as Neural Radiance Fields (NeRF) address some of these challenges, their high computational costs make them unsuitable for real-time applications. Additionally, existing 3D Gaussian Splatting (3DGS) methods often focus on photometric consistency, neglecting geometric accuracy and failing to exploit SLAM's dynamic depth and pose updates for scene refinement. We propose a framework integrating dense SLAM with 3DGS for real-time, high-fidelity dense reconstruction. Our approach introduces SLAM-Informed Adaptive Densification, which dynamically updates and densifies the Gaussian model by leveraging dense point clouds from SLAM. Additionally, we incorporate Geometry-Guided Optimization, which combines edge-aware geometric constraints and photometric consistency to jointly optimize the appearance and geometry of the 3DGS scene representation, enabling detailed and accurate SLAM mapping reconstruction. Experiments on the Replica and TUM-RGBD datasets demonstrate the effectiveness of our approach, achieving state-of-the-art results among monocular systems. Specifically, our method achieves a PSNR of 36.864, SSIM of 0.985, and LPIPS of 0.040 on Replica, representing improvements of 10.7%, 6.4%, and 49.4%, respectively, over the previous SOTA. On TUM-RGBD, our method outperforms the closest baseline by 10.2%, 6.6%, and 34.7% in the same metrics. These results highlight the potential of our framework in bridging the gap between photometric and geometric dense 3D scene representations, paving the way for practical and efficient monocular dense reconstruction.

三维重建SLAM高斯喷溅实时建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。