arXiv:2506.05965cs.CV2025-06ICRA被引 17

首个基于单目RGB的动态场景3D高斯溅射SLAM,有效抑制运动物体干扰。

Dy3DGS-SLAM: Monocular 3D Gaussian Splatting SLAM for Dynamic Environments

  • 融合光流与深度掩码的概率模型生成动态掩码,约束跟踪尺度
  • 单次网络迭代即提升渲染几何精度,动态像素损失减少遮挡干扰
  • 适用于无深度传感器的单目系统,适合真实动态环境建图

基于神经辐射场(NeRF)或3D高斯溅射(3DGS)的当前同步定位与建图(SLAM)方法在静态场景重建中表现优异,但在包含移动物体的真实动态环境中仍面临跟踪与重建难题。现有基于NeRF的动态SLAM多依赖RGB-D输入,极少支持纯单目RGB输入。为此,我们提出Dy3DGS-SLAM,首个采用单目RGB输入的动态场景3D高斯溅射SLAM方法。为应对动态干扰,我们通过概率模型融合光流掩码与深度掩码,生成融合动态掩码;仅需一次网络迭代即可约束跟踪尺度并优化渲染几何。基于该掩码,设计新型运动损失以约束位姿估计网络。在建图阶段,利用动态像素的渲染损失、颜色与深度信息,消除动态物体带来的瞬时干扰与遮挡。实验表明,Dy3DGS-SLAM在动态环境中的跟踪与渲染性能达到领先水平,优于或媲美现有RGB-D方法。

原文摘要 · Abstract (English)

Current Simultaneous Localization and Mapping (SLAM) methods based on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting excel in reconstructing static 3D scenes but struggle with tracking and reconstruction in dynamic environments, such as real-world scenes with moving elements. Existing NeRF-based SLAM approaches addressing dynamic challenges typically rely on RGB-D inputs, with few methods accommodating pure RGB input. To overcome these limitations, we propose Dy3DGS-SLAM, the first 3D Gaussian Splatting (3DGS) SLAM method for dynamic scenes using monocular RGB input. To address dynamic interference, we fuse optical flow masks and depth masks through a probabilistic model to obtain a fused dynamic mask. With only a single network iteration, this can constrain tracking scales and refine rendered geometry. Based on the fused dynamic mask, we designed a novel motion loss to constrain the pose estimation network for tracking. In mapping, we use the rendering loss of dynamic pixels, color, and depth to eliminate transient interference and occlusion caused by dynamic objects. Experimental results demonstrate that Dy3DGS-SLAM achieves state-of-the-art tracking and rendering in dynamic environments, outperforming or matching existing RGB-D methods.

SLAM3D高斯动态建图单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。