针对内窥镜图像优化深度估计,提升手术精度与安全。
Advancing Depth Anything Model for Unsupervised Monocular Depth Estimation in Endoscopy
- 用低秩适配+深度可分离残差块改进模型跨尺度适应性。
- 在SCARED和Hamlyn数据集上达到顶尖性能,参数量极低。
- 适合需要轻量化高精度深度估计的微创手术场景。
深度估计是三维重建的核心,在微创内窥镜手术中至关重要。然而,现有深度估计网络多依赖传统卷积神经网络,难以捕捉全局信息。基础模型虽具潜力,但多数基于自然图像训练,应用于内窥镜图像时表现不佳。本文提出一种针对Depth Anything模型的新微调策略,结合基于内在信息的无监督单目深度估计框架。采用基于随机向量的低秩适配技术,增强模型对不同尺度的适应能力;并设计基于深度可分离卷积的残差块,弥补变换器在局部特征捕捉上的不足。在SCARED和Hamlyn数据集上的实验表明,该方法在保持极少可训练参数的同时,实现最先进性能。应用于微创内窥镜手术可显著提升外科医生的空间感知能力,从而提高操作精度与安全性。
原文摘要 · Abstract (English)
Depth estimation is a cornerstone of 3D reconstruction and plays a vital role in minimally invasive endoscopic surgeries. However, most current depth estimation networks rely on traditional convolutional neural networks, which are limited in their ability to capture global information. Foundation models offer a promising approach to enhance depth estimation, but those models currently available are primarily trained on natural images, leading to suboptimal performance when applied to endoscopic images. In this work, we introduce a novel fine-tuning strategy for the Depth Anything Model and integrate it with an intrinsic-based unsupervised monocular depth estimation framework. Our approach includes a low-rank adaptation technique based on random vectors, which improves the model's adaptability to different scales. Additionally, we propose a residual block built on depthwise separable convolution to compensate for the transformer's limited ability to capture local features. Our experimental results on the SCARED dataset and Hamlyn dataset show that our method achieves state-of-the-art performance while minimizing the number of trainable parameters. Applying this method in minimally invasive endoscopic surgery can enhance surgeons' spatial awareness, thereby improving the precision and safety of the procedures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。