arXiv:2509.18642cs.CV2025-09中稿 · MICCAI 2025 DEMI W…被引 1

构建首个内镜图像无监督度量深度基准与合成数据集,提升模型泛化能力。

Zero-shot Monocular Metric Depth for Endoscopic Images

  • 提出合成数据集EndoSynth,含真实度量深度与分割掩码。
  • 用合成数据微调后,模型在真实内镜图像上精度显著提升。
  • 为临床场景下的内镜深度估计提供可复现的基准与资源。

近年来,由于基础模型和基于Transformer的网络的快速发展,单目相对与度量深度估计取得了显著进展。当这些方法开始应用于内镜图像领域时,仍缺乏稳健的基准和高质量数据集。本文通过在真实、未见过的内镜图像上评估最先进的(度量与相对)深度估计模型,构建了一个全面的基准,提供了关于其在临床场景中泛化能力与性能的关键洞察。此外,我们引入并发布了全新的合成数据集EndoSynth,包含内镜手术器械及其真实度量深度和分割掩码,旨在弥合合成与真实数据之间的差距。实验表明,使用该合成数据集对深度基础模型进行微调,可显著提升其在多数未见真实数据上的精度。本工作通过提供基准与合成数据集,推动了内镜图像深度估计的发展,为未来研究提供了重要资源。项目页面、EndoSynth数据集及训练权重已公开于 https://github.com/TouchSurgery/EndoSynth。

原文摘要 · Abstract (English)

Monocular relative and metric depth estimation has seen a tremendous boost in the last few years due to the sharp advancements in foundation models and in particular transformer based networks. As we start to see applications to the domain of endoscopic images, there is still a lack of robust benchmarks and high-quality datasets in that area. This paper addresses these limitations by presenting a comprehensive benchmark of state-of-the-art (metric and relative) depth estimation models evaluated on real, unseen endoscopic images, providing critical insights into their generalisation and performance in clinical scenarios. Additionally, we introduce and publish a novel synthetic dataset (EndoSynth) of endoscopic surgical instruments paired with ground truth metric depth and segmentation masks, designed to bridge the gap between synthetic and real-world data. We demonstrate that fine-tuning depth foundation models using our synthetic dataset boosts accuracy on most unseen real data by a significant margin. By providing both a benchmark and a synthetic dataset, this work advances the field of depth estimation for endoscopic images and serves as an important resource for future research. Project page, EndoSynth dataset and trained weights are available at https://github.com/TouchSurgery/EndoSynth.

内镜深度合成数据零样本医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。