解决内镜图像纹理弱、光照变化大时的深度估计难题
Occlusion-Aware Self-Supervised Monocular Depth Estimation for Weak-Texture Endoscopic Images
- 用遮挡掩码模拟视角遮挡,增强模型对部分可见场景的鲁棒性
- 在无纹理区域通过非负矩阵分解聚类激活值生成伪标签,提升分割精度
- 在多个数据集上表现优于现有方法,适合真实内镜场景应用
我们提出一种面向内镜场景的自监督单目深度估计网络,旨在从单张图像中推断胃肠道内部深度。现有方法虽准确,但通常假设光照一致,而实际中胃肠蠕动导致动态光照和遮挡,破坏几何判断并降低自监督信号可靠性,影响深度重建质量。为此,我们设计了一种遮挡感知的自监督框架:首先引入遮挡掩码进行数据增强,通过模拟视点相关的遮挡场景生成伪标签,提升模型在部分遮挡下的深度特征学习能力;其次,利用非负矩阵分解引导语义分割,对卷积激活进行聚类以生成无纹理区域的伪标签,从而改善分割准确性并缓解光照变化带来的信息丢失。在SCARED数据集上的实验表明,本方法在自监督深度估计中达到当前最优性能;在Endo-SLAM和SERV-CT数据集上的评估也证明其在多种内镜环境中的强泛化能力。
原文摘要 · Abstract (English)
We propose a self-supervised monocular depth estimation network tailored for endoscopic scenes, aiming to infer depth within the gastrointestinal tract from monocular images. Existing methods, though accurate, typically assume consistent illumination, which is often violated due to dynamic lighting and occlusions caused by GI motility. These variations lead to incorrect geometric interpretations and unreliable self-supervised signals, degrading depth reconstruction quality. To address this, we introduce an occlusion-aware self-supervised framework. First, we incorporate an occlusion mask for data augmentation, generating pseudo-labels by simulating viewpoint-dependent occlusion scenarios. This enhances the model's ability to learn robust depth features under partial visibility. Second, we leverage semantic segmentation guided by non-negative matrix factorization, clustering convolutional activations to generate pseudo-labels in texture-deprived regions, thereby improving segmentation accuracy and mitigating information loss from lighting changes. Experimental results on the SCARED dataset show that our method achieves state-of-the-art performance in self-supervised depth estimation. Additionally, evaluations on the Endo-SLAM and SERV-CT datasets demonstrate strong generalization across diverse endoscopic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。