arXiv:2512.10379cs.CV2025-12

用自监督方法提升内窥镜图像匹配精度,解决手术视觉难题。

Self-Supervised Contrastive Embedding Adaptation for Endoscopic Image Matching

  • 通过新视角合成生成真实对应关系,用于对比学习训练
  • 在SCARED数据集上匹配精度更高,视差误差更低
  • 适合需要高精度图像匹配的手术导航与三维重建场景

精准的空间理解对术中导航、增强现实融合和环境感知至关重要。微创手术中,视觉是唯一的术中输入方式,建立内窥镜帧间的像素级对应关系对于三维重建、相机追踪和场景理解极为关键。然而,手术领域存在弱透视线索、非朗伯组织反射及复杂可变形解剖结构,导致传统计算机视觉方法性能下降。尽管深度学习在自然场景表现优异,其特征并不天然适用于精细匹配的手术图像,需针对性适配。本研究提出一种新型深度学习流水线,用于建立内窥镜图像对的特征对应关系,并设计自监督优化框架进行模型训练。该方法利用新颖视角合成生成真实内点对应关系,用于对比学习中的三元组挖掘。通过此自监督策略,我们在DINOv2主干网络基础上增加一个额外的Transformer层,专门优化以生成可通过余弦相似度阈值直接匹配的嵌入表示。实验评估表明,该流水线在SCARED数据集上优于现有最先进方法,匹配精度更高,且极小基线误差更低。所提框架为实现更准确的高级计算机视觉应用在手术内窥镜中的落地提供了重要贡献。

原文摘要 · Abstract (English)

Accurate spatial understanding is essential for image-guided surgery, augmented reality integration and context awareness. In minimally invasive procedures, where visual input is the sole intraoperative modality, establishing precise pixel-level correspondences between endoscopic frames is critical for 3D reconstruction, camera tracking, and scene interpretation. However, the surgical domain presents distinct challenges: weak perspective cues, non-Lambertian tissue reflections, and complex, deformable anatomy degrade the performance of conventional computer vision techniques. While Deep Learning models have shown strong performance in natural scenes, their features are not inherently suited for fine-grained matching in surgical images and require targeted adaptation to meet the demands of this domain. This research presents a novel Deep Learning pipeline for establishing feature correspondences in endoscopic image pairs, alongside a self-supervised optimization framework for model training. The proposed methodology leverages a novel-view synthesis pipeline to generate ground-truth inlier correspondences, subsequently utilized for mining triplets within a contrastive learning paradigm. Through this self-supervised approach, we augment the DINOv2 backbone with an additional Transformer layer, specifically optimized to produce embeddings that facilitate direct matching through cosine similarity thresholding. Experimental evaluation demonstrates that our pipeline surpasses state-of-the-art methodologies on the SCARED datasets improved matching precision and lower epipolar error compared to the related work. The proposed framework constitutes a valuable contribution toward enabling more accurate high-level computer vision applications in surgical endoscopy.

内窥镜匹配自监督学习医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。