arXiv:2505.08178cs.CV2025-05

用单目深度指导,提升腹腔镜图像的视差估计精度

Monocular Depth Guided Occlusion-Aware Disparity Refinement via Semi-supervised Learning in Laparoscopic Images

  • 利用单目深度信息缓解遮挡影响,改进视差图
  • 在SCARED数据集上EPE降低12.3%,RMSE下降9.7%
  • 适合医疗影像中数据少、遮挡多的场景

腹腔镜图像中遮挡和标注数据稀缺是视差估计的主要挑战。本文提出深度引导的遮挡感知视差精化网络(DGORNet),通过不受遮挡影响的单目深度信息来优化视差图。引入位置嵌入模块提供显式空间上下文,增强特征定位与精化能力。同时设计光流差异损失(OFDLoss)用于无标签数据,利用视频帧间的时序连续性提升动态手术场景下的鲁棒性。在SCARED数据集上的实验表明,DGORNet在端点误差(EPE)和均方根误差(RMSE)上优于现有方法,尤其在遮挡区域和纹理缺失区域表现更优。消融实验证实位置嵌入和光流差异损失对提升空间与时间一致性的关键作用。结果表明DGORNet能有效解决腹腔镜视差估计中的数据与遮挡难题。

原文摘要 · Abstract (English)

Occlusion and the scarcity of labeled surgical data are significant challenges in disparity estimation for stereo laparoscopic images. To address these issues, this study proposes a Depth Guided Occlusion-Aware Disparity Refinement Network (DGORNet), which refines disparity maps by leveraging monocular depth information unaffected by occlusion. A Position Embedding (PE) module is introduced to provide explicit spatial context, enhancing the network's ability to localize and refine features. Furthermore, we introduce an Optical Flow Difference Loss (OFDLoss) for unlabeled data, leveraging temporal continuity across video frames to improve robustness in dynamic surgical scenes. Experiments on the SCARED dataset demonstrate that DGORNet outperforms state-of-the-art methods in terms of End-Point Error (EPE) and Root Mean Squared Error (RMSE), particularly in occlusion and texture-less regions. Ablation studies confirm the contributions of the Position Embedding and Optical Flow Difference Loss, highlighting their roles in improving spatial and temporal consistency. These results underscore DGORNet's effectiveness in enhancing disparity estimation for laparoscopic surgery, offering a practical solution to challenges in disparity estimation and data limitations.

视差估计医疗影像半监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。