arXiv:2608.00415cs.CVcs.AI2026-08中稿 · MICCAI 2026 @ The …

提升内镜场景下深度估计的泛化能力,解决光照干扰问题。

Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment

论文配图:Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment
图 1 · 摘自论文原文
  • 采用轻量级专家混合机制实现高效微调,适应不同内镜场景。
  • 引入内在图像对齐,在零样本测试中显著提升精度。
  • 适用于医疗内镜深度估计,尤其适合光照变化大的场景。

内镜手术中的深度估计对三维感知至关重要。然而,光照干扰和不同内镜场景下的特征多样性仍是泛化深度估计与自身运动估计的主要挑战。为此,提出一种新型自监督框架EndoMINI,用于内镜场景深度估计。具体而言,提出低秩专家混合(MiLoRE)以实现参数高效的微调,并增强模型对不同场景特征的适应能力。同时,引入内在图像对齐(IIA)至训练损失,通过新型内在图像分解网络缓解内镜中光反射的影响。该方法在SCARED数据集上进行有监督深度估计评估,并在两个内镜数据集Hamlyn和SERV-CT上进行零样本深度估计对比实验,结果表明所提模型性能优异,且主要贡献效果显著。

原文摘要 · Abstract (English)

Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interference and feature diversity in various endoscopic scenes are still challenges for generalizable depth estimation and ego-motion estimation. Based on this, a novel self-supervised framework, EndoMINI, is proposed for depth estimation in endoscopic scenes. Specifically, mixture of low-rank experts (MiLoRE) is proposed to perform parameter-efficient fine-tuning, which can also boost the model adaptation to scenes with different characteristics. Meanwhile, an intrinsic image alignment (IIA) is introduced into the training loss to alleviate the influence of light reflectance in endoscopy with a novel intrinsic image decomposition network. The proposed method is evaluated on SCARED datasets for supervised depth estimation, and two endoscopic datasets, Hamlyn and SERV-CT, for zero-shot depth estimation, compared with state-of-the-art works as well. The experimental results demonstrate outstanding performance of the proposed model and the effects of the main contributions.

深度估计内镜视觉自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。