提出基于3D高斯点云的内镜非刚性定位与建图方法,有效分离运动与形变干扰。
NRGS-SLAM: Monocular Non-Rigid SLAM for Endoscopy via Deformation-Aware 3D Gaussian Splatting
- 用可学习形变概率的3D高斯表示,通过贝叶斯自监督优化解耦形变与相机运动
- 在多个公开数据集上实现姿态误差降低50%、重建质量显著提升
- 适合需要高精度内镜导航与三维重建的医疗场景应用
视觉同步定位与建图(V-SLAM)是自主感知与导航的基础能力。然而,内镜场景中软组织持续形变破坏了刚性假设,导致相机自身运动与内部形变之间存在强烈耦合模糊性。尽管近期单目非刚性SLAM方法取得进展,但仍缺乏有效解耦机制,且依赖稀疏或低保真场景表示,易引发跟踪漂移和重建质量差的问题。为此,我们提出NRGS-SLAM,一种基于3D高斯点云的单目内镜非刚性SLAM系统。为解决耦合模糊性,引入形变感知的3D高斯地图,每个高斯原语附加可学习形变概率,通过贝叶斯自监督策略优化,无需外部非刚性标签。基于此表示,设计可变形跟踪模块,优先关注低形变区域,实现鲁棒的粗到精位姿估计,并高效完成逐帧形变更新。进一步设计可变形映射模块,逐步扩展并优化地图,在表示能力与计算效率间取得平衡。此外,统一的鲁棒几何损失融合外部几何先验,缓解单目非刚性SLAM固有的病态性。在多个公开内镜数据集上的大量实验表明,NRGS-SLAM相比现有最优方法,相机位姿估计误差最多降低50%(RMSE),并生成更高质量的逼真重建结果。全面的消融实验验证了关键设计的有效性。代码将在论文录用后公开。
原文摘要 · Abstract (English)
Visual simultaneous localization and mapping (V-SLAM) is a fundamental capability for autonomous perception and navigation. However, endoscopic scenes violate the rigidity assumption due to persistent soft-tissue deformations, creating a strong coupling ambiguity between camera ego-motion and intrinsic deformation. Although recent monocular non-rigid SLAM methods have made notable progress, they often lack effective decoupling mechanisms and rely on sparse or low-fidelity scene representations, which leads to tracking drift and limited reconstruction quality. To address these limitations, we propose NRGS-SLAM, a monocular non-rigid SLAM system for endoscopy based on 3D Gaussian Splatting. To resolve the coupling ambiguity, we introduce a deformation-aware 3D Gaussian map that augments each Gaussian primitive with a learnable deformation probability, optimized via a Bayesian self-supervision strategy without requiring external non-rigidity labels. Building on this representation, we design a deformable tracking module that performs robust coarse-to-fine pose estimation by prioritizing low-deformation regions, followed by efficient per-frame deformation updates. A carefully designed deformable mapping module progressively expands and refines the map, balancing representational capacity and computational efficiency. In addition, a unified robust geometric loss incorporates external geometric priors to mitigate the inherent ill-posedness of monocular non-rigid SLAM. Extensive experiments on multiple public endoscopic datasets demonstrate that NRGS-SLAM achieves more accurate camera pose estimation (up to 50\% reduction in RMSE) and higher-quality photo-realistic reconstructions than state-of-the-art methods. Comprehensive ablation studies further validate the effectiveness of our key design choices. Source code will be publicly available upon paper acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。