用主动学习减少手术视频标注工作量,仅需一半人工即可完成精准分割。
Active Learning for Efficient Annotation of Surgical Videos with Weak Supervision

- 结合主动学习与双损失优化,通过迭代生成伪掩码引导专家修正。
- 训练结束时标注效率提升50%,相比全手动标注大幅减少工作量。
- 适合需要高效构建手术工具分割模型的研究者和临床应用团队。
腹腔镜视频的时空精确标注耗时且需专业知识。我们提出一种人机协同的知识获取框架,融合主动学习与双损失优化,显著降低自动定位与分割手术区域所需的标注成本。该方法利用基础模型从视频生成时间一致的类别激活图(CAM),采用两种互补的训练目标:基于视频级工具存在标签的弱监督损失,以及基于主动学习获取的人工校正标注的图像级掩码损失。无需预先进行密集像素级标注,该流程通过迭代生成伪掩码,指导专家对模型已捕捉知识进行精炼。实验表明,相较于完全手动标注,本框架在训练结束时将标注工作量减少50%。通过避免初始阶段依赖大规模全标注数据集,该框架支持手术工具分割模型的可扩展开发。这种迭代式人机协同机制以最小专家投入实现高效知识积累,为扩展至更大、更多样化的数据集及真实临床场景提供了实用可行的策略。
原文摘要 · Abstract (English)
Precise spatial-temporal annotation of laparoscopic videos is time-consuming and requires expert knowledge. We propose a human-in-the-loop knowledge acquisition framework that combines active learning with dual-loss optimization to significantly reduce the annotation effort needed for automatic localization and segmentation of objects in the surgical field. Our method employs a foundation model to generate temporally consistent class activation maps (CAMs) from video using two complementary training objectives: a weak supervision loss on video-level tool presence labels for weakly annotated data, and an image-level mask loss on human-corrected annotations obtained through active learning. Rather than requiring dense pixel-level annotation upfront, our pipeline iteratively proposes pseudo-masks that guide the expert annotator to refine the knowledge previously captured by the model. We demonstrate that our framework reduces the effort of surgical video annotation by 50% by the end of training in comparison to fully manual annotation. Through eliminating the need for large, fully annotated datasets from the start, this framework enables scalability to the development of surgical tool segmentation models. This iterative human-in-the-loop refinement supports efficient knowledge acquisition with minimal expert input, providing a practical and deployable strategy for expanding tool segmentation to larger, more diverse datasets and real-world clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。