针对内窥镜场景多样性,提出可泛化的单目深度估计算法
EndoGMDE: Generalizable Monocular Depth Estimation with Mixture of Low-Rank Experts for Diverse Endoscopic Scenes
- 采用分块动态低秩专家混合机制,自适应选择少量可训练参数的专家
- 在SCARED和SimCol数据集上优于现有方法,零样本迁移在多个数据集表现最佳
- 适合需要高精度3D重建与手术导航的微创医疗场景
自监督单目深度估计是内窥镜中低成本高效三维感知与测量的重要任务。然而,光照条件多样性和场景特征差异仍是内窥镜深度估计的主要挑战。本文提出一种新颖的自监督框架,用于多样化内窥镜场景下的单目深度估计。首先,针对不同组织带来的多样化场景特征,提出一种基于块的动态低秩专家混合模块,以高效微调基础模型。该模块根据输入特征自适应选择少量可训练参数的专家进行加权推理,各专家基于块级泛化能力分配。此外,设计了一种新型自监督训练框架,联合处理亮度不一致与反射干扰问题。所提方法在SCARED和SimCol数据集上超越当前最优方法,且在C3VD、Hamlyn和SERV-CT数据集上实现最佳零样本泛化性能。模型在三维重建与自身运动估计任务中表现优异,可助力微创测量与手术中的精确内窥镜应用。评估代码将在录用后公开,演示视频见:https://endo-gmde.netlify.app/。
原文摘要 · Abstract (English)
Self-supervised monocular depth estimation is a significant task for low-cost and efficient 3D scene perception and measurement in endoscopy. However, the variety of illumination conditions and scene features is still the primary challenges for depth estimation in endoscopic scenes. In this work, a novel self-supervised framework is proposed for monocular depth estimation in diverse endoscopy. Firstly, considering the diverse features in endoscopic scenes with different tissues, a novel block-wise mixture of dynamic low-rank experts is proposed to efficiently finetune the foundation model for endoscopic depth estimation. In the proposed module, based on the input feature, different experts with a small amount of trainable parameters are adaptively selected for weighted inference, from low-rank experts which are allocated based on the generalization of each block. Moreover, a novel self-supervised training framework is proposed to jointly cope with brightness inconsistency and reflectance interference. The proposed method outperforms state-of-the-art works on SCARED dataset and SimCol dataset. Furthermore, the proposed network also achieves the best generalization based on zero-shot depth estimation on C3VD, Hamlyn and SERV-CT dataset. The outstanding performance of our model is further demonstrated with 3D reconstruction and ego-motion estimation. The proposed method could contribute to accurate endoscopy for minimally invasive measurement and surgery. The evaluation codes will be released upon acceptance, while the demo videos can be found on: https://endo-gmde.netlify.app/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。