arXiv:2508.17916cs.CV2025-08被引 2

用双基础模型提升内窥镜单目深度估计,解决光照与纹理复杂问题。

EndoUFM: Utilizing Foundation Models for Monocular depth estimation of endoscopic images

  • 引入双基础模型+自适应微调,增强对内窥镜图像的语义理解。
  • 在4个数据集上达到最优性能,参数量小且推理快。
  • 适合需要高精度空间感知的微创手术导航系统研发者。

深度估计是微创内窥镜手术中三维重建的关键环节。然而,现有单目深度估计方法在手术环境多变光照与复杂纹理下表现有限。尽管基础模型有望提升性能,但其预训练自然图像与目标内窥镜图像间的领域差异导致显著语义感知缺失。本文提出无监督单目深度估计框架EndoUFM,创新性地利用双基础模型,通过强大的先验知识提升深度估计效果。该框架采用新型自适应微调策略,结合随机向量低秩适配(RVLoRA)以增强模型适应性,并设计基于深度可分离卷积的残差模块(Res-DSC)以更好捕捉细粒度局部特征。此外,引入掩码引导平滑损失,强化解剖结构内深度的一致性。在SCARED、Hamlyn、SERV-CT和EndoNeRF四个数据集上的大量实验表明,本方法在保持高效模型规模的同时达到当前最优性能。该工作有助于提升外科医生术中空间感知能力,增强手术精度与安全性,对增强现实与导航系统具有重要意义。代码已公开于https://github.com/RealMindyY/EndoUFM。

原文摘要 · Abstract (English)

Depth estimation is a foundational component for 3D reconstruction in minimally invasive endoscopic surgeries. However, existing monocular depth estimation techniques often exhibit limited performance to the varying illumination and complex textures of the surgical environment. While applying foundation models offers a promising approach to enhance the depth estimation performance, the domain gap between the natural images used for pre-training and the target endoscopic images leads to significant semantic perception deficiencies. In this study, EndoUFM is introduced as an unsupervised monocular depth estimation framework that innovatively \underline{U}tilizes dual Foundation Models for Endoscopic images, thereby enhancing the depth estimation performance by leveraging the powerful pre-learned priors. The framework features a novel adaptive fine-tuning strategy that incorporates Random Vector Low-Rank Adaptation (RVLoRA) to enhance model adaptability, and a Residual block based on Depthwise Separable Convolution (Res-DSC) to improve the capture of fine-grained local features. A mask-guided smoothness loss is also introduced to enforce depth consistency within anatomical structures. Extensive experiments on the SCARED, Hamlyn, SERV-CT, and EndoNeRF datasets confirm that our method achieves state-of-the-art performance while maintaining an efficient model size. This work contributes to augmenting surgeons' spatial perception during minimally invasive procedures, thereby enhancing surgical precision and safety, with crucial implications for augmented reality and navigation systems. Our code is available at https://github.com/RealMindyY/EndoUFM.

深度估计内窥镜基础模型微创手术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。