通过不确定性融合提升内窥镜深度估计泛化能力
Generalizing monocular colonoscopy image depth estimation by uncertainty-based global and local fusion network
- 结合CNN局部特征与Transformer全局信息,用不确定性机制融合
- 在多个数据集上实现无微调直接跨域泛化,误差低于1.25mm
- 适合临床内窥镜导航、息肉检测等实际应用
深度估计对内窥镜导航与操作至关重要,但在结肠等真实临床场景中获取真值深度图极具挑战。本研究提出一种融合卷积神经网络(CNN)捕捉局部特征与Transformer捕捉全局信息的框架,并设计基于不确定性的融合模块,以识别两者互补贡献。该网络可仅使用仿真数据训练,无需任何微调即可直接应用于未见的临床数据。实验验证了方法在多个数据集上的优异泛化性能,涵盖不同解剖结构。真实临床场景的定性分析也证实了其鲁棒性。结果表明,通过CNN-Transformer架构与不确定性融合块的结合,显著提升了模拟与真实内窥镜环境中的深度估计性能与泛化能力。本研究为复杂临床条件下内窥镜深度图估计提供了新范式,可支撑自动导航、息肉检测与分割等任务。
原文摘要 · Abstract (English)
Objective: Depth estimation is crucial for endoscopic navigation and manipulation, but obtaining ground-truth depth maps in real clinical scenarios, such as the colon, is challenging. This study aims to develop a robust framework that generalizes well to real colonoscopy images, overcoming challenges like non-Lambertian surface reflection and diverse data distributions. Methods: We propose a framework combining a convolutional neural network (CNN) for capturing local features and a Transformer for capturing global information. An uncertainty-based fusion block was designed to enhance generalization by identifying complementary contributions from the CNN and Transformer branches. The network can be trained with simulated datasets and generalize directly to unseen clinical data without any fine-tuning. Results: Our method is validated on multiple datasets and demonstrates an excellent generalization ability across various datasets and anatomical structures. Furthermore, qualitative analysis in real clinical scenarios confirmed the robustness of the proposed method. Conclusion: The integration of local and global features through the CNN-Transformer architecture, along with the uncertainty-based fusion block, improves depth estimation performance and generalization in both simulated and real-world endoscopic environments. Significance: This study offers a novel approach to estimate depth maps for endoscopy images despite the complex conditions in clinic, serving as a foundation for endoscopic automatic navigation and other clinical tasks, such as polyp detection and segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。