arXiv:2409.07801cs.CV2024-09被引 2

用自监督方法在极少标注下定位手术视频中的工具与结构

SURGIVID: Annotation-Efficient Surgical Video Object Discovery

  • 通过自监督发现手术视频中显著的工具和解剖结构
  • 仅需36个标注即达到全监督模型性能
  • 适合医疗数据少、标注难的手术分析场景

手术场景包含手术质量的关键信息。对工具和解剖结构进行像素级定位是实现显微或内窥镜手术视图深度分析的第一步。传统方法依赖全监督训练,标注成本高且常需医学专业知识。鉴于标准化手术流程产生的大量手术视频,我们提出一种标注高效的手术场景语义分割框架。采用基于图像的自监督对象发现方法,识别手术视频中最显著的工具和解剖结构,并通过最小监督微调进一步优化。仅使用36个标注标签的无监督设置即可达到与全监督分割模型相当的定位性能。此外,利用手术阶段标签作为弱标签可更好引导模型关注手术工具,使工具定位性能提升约2%。在CaDIS数据集上的充分消融实验验证了该方案在极小或无监督条件下发现相关手术对象的有效性。

原文摘要 · Abstract (English)

Surgical scenes convey crucial information about the quality of surgery. Pixel-wise localization of tools and anatomical structures is the first task towards deeper surgical analysis for microscopic or endoscopic surgical views. This is typically done via fully-supervised methods which are annotation greedy and in several cases, demanding medical expertise. Considering the profusion of surgical videos obtained through standardized surgical workflows, we propose an annotation-efficient framework for the semantic segmentation of surgical scenes. We employ image-based self-supervised object discovery to identify the most salient tools and anatomical structures in surgical videos. These proposals are further refined within a minimally supervised fine-tuning step. Our unsupervised setup reinforced with only 36 annotation labels indicates comparable localization performance with fully-supervised segmentation models. Further, leveraging surgical phase labels as weak labels can better guide model attention towards surgical tools, leading to $\sim 2\%$ improvement in tool localization. Extensive ablation studies on the CaDIS dataset validate the effectiveness of our proposed solution in discovering relevant surgical objects with minimal or no supervision.

手术视频自监督语义分割少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。