融合深度感知与视觉先验,提升单目内窥镜定位与组织重建精度。
Unifying Scale-Aware Depth Prediction and Perceptual Priors for Monocular Endoscope Pose Estimation and Tissue Reconstruction
- 引入MAPIS-Depth模块,结合深度模型与优化算法生成伪度量深度
- 通过时序感知的像素匹配与感知相似性融合,减少组织变形带来的伪影
- 适用于微创手术导航,尤其适合纹理少、视角受限场景
准确的内窥镜位姿估计与三维组织表面重建能显著提升单目微创手术的导航精度与空间感知能力。然而,单目内窥镜面临深度模糊、生理组织形变、运动不一致、纹理保真度低及视野受限等挑战。为此,提出一种统一框架,集成尺度感知深度预测与时序约束的感知精修。该框架采用新型MAPIS-Depth模块,利用Depth Pro进行鲁棒初始化,结合Depth Anything实现高效逐帧深度预测,并通过L-BFGS-B优化生成伪度量深度。这些深度估计通过RAFT计算像素对应关系,基于LPIPS感知相似性自适应融合光流扭曲图像,有效降低组织形变和运动引起的伪影。为确保伪RGBD帧的准确配准,引入WEMA-RTDL模块,联合优化旋转与平移。最后,采用截断有符号距离函数的体素融合与Marching Cubes算法提取完整3D表面网格。在HEVD和SCARED数据集上的实验表明,该方法在消融分析与对比测试中均优于现有最优方法。
原文摘要 · Abstract (English)
Accurate endoscope pose estimation and 3D tissue surface reconstruction significantly enhances monocular minimally invasive surgical procedures by enabling accurate navigation and improved spatial awareness. However, monocular endoscope pose estimation and tissue reconstruction face persistent challenges, including depth ambiguity, physiological tissue deformation, inconsistent endoscope motion, limited texture fidelity, and a restricted field of view. To overcome these limitations, a unified framework for monocular endoscopic tissue reconstruction that integrates scale-aware depth prediction with temporally-constrained perceptual refinement is presented. This framework incorporates a novel MAPIS-Depth module, which leverages Depth Pro for robust initialisation and Depth Anything for efficient per-frame depth prediction, in conjunction with L-BFGS-B optimisation, to generate pseudo-metric depth estimates. These estimates are temporally refined by computing pixel correspondences using RAFT and adaptively blending flow-warped frames based on LPIPS perceptual similarity, thereby reducing artefacts arising from physiological tissue deformation and motion. To ensure accurate registration of the synthesised pseudo-RGBD frames from MAPIS-Depth, a novel WEMA-RTDL module is integrated, optimising both rotation and translation. Finally, truncated signed distance function-based volumetric fusion and marching cubes are applied to extract a comprehensive 3D surface mesh. Evaluations on HEVD and SCARED, with ablation and comparative analyses, demonstrate the framework's robustness and superiority over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。