arXiv:2506.01471cs.CV2025-06中稿 · MICCAI 2025被引 4

用少量标注数据实现精准手术阶段识别,提升效率与效果。

SemiVT-Surge: Semi-Supervised Video Transformer for Surgical Phase Recognition

  • 基于视频变换器与伪标签框架,结合时序一致性与对比学习。
  • 在RAMIE上提升4.9%准确率,仅用1/4标注数据达全监督效果。
  • 适合医疗视频分析、资源受限场景下的手术阶段识别研究。

精准的手术阶段识别对计算机辅助干预和手术视频分析至关重要。标注长段手术视频耗时耗力,促使研究转向利用未标注数据以少量标注实现优异性能。尽管自监督学习通过大规模预训练加小规模标注微调获得关注,半监督方法在手术领域仍鲜有探索。本文提出一种基于视频变换器的模型,配备稳健的伪标签框架,融合未标注数据的时序一致性正则化与基于类别原型的对比学习,同时利用标注数据和伪标签优化特征空间。在私有RAMIE(机器人辅助微创食管切除术)数据集和公开Cholec80数据集上进行大量实验,结果表明:引入未标注数据后,在RAMIE上准确率提升4.9%,在Cholec80上仅使用1/4标注数据即达到接近全监督的性能。研究为半监督手术阶段识别建立强基准,推动该领域未来发展。

原文摘要 · Abstract (English)

Accurate surgical phase recognition is crucial for computer-assisted interventions and surgical video analysis. Annotating long surgical videos is labor-intensive, driving research toward leveraging unlabeled data for strong performance with minimal annotations. Although self-supervised learning has gained popularity by enabling large-scale pretraining followed by fine-tuning on small labeled subsets, semi-supervised approaches remain largely underexplored in the surgical domain. In this work, we propose a video transformer-based model with a robust pseudo-labeling framework. Our method incorporates temporal consistency regularization for unlabeled data and contrastive learning with class prototypes, which leverages both labeled data and pseudo-labels to refine the feature space. Through extensive experiments on the private RAMIE (Robot-Assisted Minimally Invasive Esophagectomy) dataset and the public Cholec80 dataset, we demonstrate the effectiveness of our approach. By incorporating unlabeled data, we achieve state-of-the-art performance on RAMIE with a 4.9% accuracy increase and obtain comparable results to full supervision while using only 1/4 of the labeled data on Cholec80. Our findings establish a strong benchmark for semi-supervised surgical phase recognition, paving the way for future research in this domain.

手术识别半监督视频变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。