arXiv:2511.02247cs.CV2025-11被引 4

通过特征对齐提升内窥镜绝对深度估计精度

Monocular absolute depth estimation from endoscopy via domain-invariant feature learning and latent consistency

  • 利用对抗学习与方向一致性对齐真实与合成图像的潜在特征
  • 在气道幻影数据上实现优于现有方法的绝对与相对深度指标
  • 不依赖图像风格迁移,适配多种骨干网络和预训练权重

单目深度估计(MDE)对指导自主医疗机器人至关重要。然而,在手术场景中从内窥镜相机获取绝对(度量)深度仍具挑战性,限制了真实内窥镜图像上的监督学习。现有图像级无监督域适应方法将带有已知深度图的合成图像转换为真实内窥镜帧的风格,并使用这些转换后的图像及其对应深度图训练深度网络。然而,真实与转换合成图像之间常存在域差距。本文提出一种潜在特征对齐方法,通过减少中心气道内窥镜视频中的域差距,提升绝对深度估计性能。该方法与图像翻译过程无关,专注于深度估计本身。具体而言,深度网络输入经翻译的合成与真实内窥镜帧,通过对抗学习与方向特征一致性学习潜在域不变特征。评估在带有手动对齐绝对深度图的中心气道幻影内窥镜视频上进行。相比当前最先进方法,本方案在绝对与相对深度指标上均取得更优表现,且在不同骨干网络与预训练权重下均一致提升。代码已公开于 https://github.com/MedICL-VU/MDE。

原文摘要 · Abstract (English)

Monocular depth estimation (MDE) is a critical task to guide autonomous medical robots. However, obtaining absolute (metric) depth from an endoscopy camera in surgical scenes is difficult, which limits supervised learning of depth on real endoscopic images. Current image-level unsupervised domain adaptation methods translate synthetic images with known depth maps into the style of real endoscopic frames and train depth networks using these translated images with their corresponding depth maps. However a domain gap often remains between real and translated synthetic images. In this paper, we present a latent feature alignment method to improve absolute depth estimation by reducing this domain gap in the context of endoscopic videos of the central airway. Our methods are agnostic to the image translation process and focus on the depth estimation itself. Specifically, the depth network takes translated synthetic and real endoscopic frames as input and learns latent domain-invariant features via adversarial learning and directional feature consistency. The evaluation is conducted on endoscopic videos of central airway phantoms with manually aligned absolute depth maps. Compared to state-of-the-art MDE methods, our approach achieves superior performance on both absolute and relative depth metrics, and consistently improves results across various backbones and pretrained weights. Our code is available at https://github.com/MedICL-VU/MDE.

深度估计内窥镜域适应医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。