arXiv:2410.07434cs.CV2024-10被引 14

微调深度模型,让手术场景下的深度估计更准

Surgical Depth Anything: Depth Estimation for Surgical Scenes using Foundation Models

  • 用手术数据微调深度模型,适应术中模糊、出血等特殊场景
  • 显著降低模糊与反光导致的深度误差,提升精度
  • 适合需要高精度3D重建的医疗影像研究者

单目深度估计对跟踪和重建算法至关重要,尤其在手术视频中。然而,术中难以获取真实深度图,使监督学习方法不可行。尽管基于运动结构(SfM)的自监督方法表现良好,但其依赖高质量相机运动,且需逐患者优化。通过利用当前先进的深度基础模型 Depth Anything,可缓解上述问题。然而,直接应用于手术场景时,该模型易受模糊、出血和反光影响,表现不佳。本文针对手术领域对 Depth Anything 进行微调,旨在生成更准确、像素级的深度图,以满足手术环境的独特需求。微调后模型显著改善了模糊与反光带来的误差,实现更可靠、精确的深度估计。

原文摘要 · Abstract (English)

Monocular depth estimation is crucial for tracking and reconstruction algorithms, particularly in the context of surgical videos. However, the inherent challenges in directly obtaining ground truth depth maps during surgery render supervised learning approaches impractical. While many self-supervised methods based on Structure from Motion (SfM) have shown promising results, they rely heavily on high-quality camera motion and require optimization on a per-patient basis. These limitations can be mitigated by leveraging the current state-of-the-art foundational model for depth estimation, Depth Anything. However, when directly applied to surgical scenes, Depth Anything struggles with issues such as blurring, bleeding, and reflections, resulting in suboptimal performance. This paper presents a fine-tuning of the Depth Anything model specifically for the surgical domain, aiming to deliver more accurate pixel-wise depth maps tailored to the unique requirements and challenges of surgical environments. Our fine-tuning approach significantly improves the model's performance in surgical scenes, reducing errors related to blurring and reflections, and achieving a more reliable and precise depth estimation.

深度估计手术视觉模型微调医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。