arXiv:2602.17785cs.CV2026-02中稿 · IPCAI2026

用边缘和光照分解提升内窥镜深度与位姿估计,无需真实标注。

Multi-Modal Monocular Endoscopic Depth and Pose Estimation with Edge-Guided Self-Supervision

  • 基于边缘图和光照解耦的自监督学习框架
  • 在真实与合成数据上均达当前最佳性能
  • 适合内窥镜导航、医学图像分析研究者

单目深度与位姿估计在结肠镜辅助导航中至关重要,可减少盲区、降低漏诊与检查不完整风险。但因缺乏纹理表面、复杂光照、形变及真实体内标注数据,该任务仍具挑战。本文提出PRISM(Pose-Refinement with Intrinsic Shading and edge Maps)框架,利用解剖结构与光照先验指导几何学习。通过学习型边缘检测器(如DexiNed或HED)提取细长高频边界,结合内在分解模块分离阴影与反射率,使模型有效利用阴影信息进行深度估计。在多个真实与合成数据集上的实验表明性能领先。消融研究揭示:(1)在真实数据上自监督训练优于在逼真假体数据上的监督训练,凸显领域真实性的优势;(2)视频帧率对模型性能影响极大,需根据数据集特性采样帧以生成高质量训练数据。

原文摘要 · Abstract (English)

Monocular depth and pose estimation play an important role in the development of colonoscopy-assisted navigation, as they enable improved screening by reducing blind spots, minimizing the risk of missed or recurrent lesions, and lowering the likelihood of incomplete examinations. However, this task remains challenging due to the presence of texture-less surfaces, complex illumination patterns, deformation, and a lack of in-vivo datasets with reliable ground truth. In this paper, we propose **PRISM** (Pose-Refinement with Intrinsic Shading and edge Maps), a self-supervised learning framework that leverages anatomical and illumination priors to guide geometric learning. Our approach uniquely incorporates edge detection and luminance decoupling for structural guidance. Specifically, edge maps are derived using a learning-based edge detector (e.g., DexiNed or HED) trained to capture thin and high-frequency boundaries, while luminance decoupling is obtained through an intrinsic decomposition module that separates shading and reflectance, enabling the model to exploit shading cues for depth estimation. Experimental results on multiple real and synthetic datasets demonstrate state-of-the-art performance. We further conduct a thorough ablation study on training data selection to establish best practices for pose and depth estimation in colonoscopy. This analysis yields two practical insights: (1) self-supervised training on real-world data outperforms supervised training on realistic phantom data, underscoring the superiority of domain realism over ground truth availability; and (2) video frame rate is an extremely important factor for model performance, where dataset-specific video frame sampling is necessary for generating high quality training data.

内窥镜自监督深度估计边缘引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。