针对白内障手术视频,提出按步骤选择整段视频的主动学习方法。
StepAL: Step-aware Active Learning for Cataract Surgical Videos
- 基于伪标签和熵加权聚类,选择不确定且步骤多样的完整视频。
- 在两个数据集上用更少标注视频达到更高识别准确率。
- 适合需要高效标注长手术视频的研究者或医疗AI开发者。
主动学习(AL)可在保持模型性能的前提下降低手术视频分析的标注成本。然而,传统AL方法针对图像或短视频设计,难以适用于长而未剪裁的手术视频中的步骤识别任务,因其存在步骤间依赖关系。这些方法通常仅选择单个帧或片段进行标注,而手术视频的标注需依赖整体上下文,因此效果不佳。为此,我们提出StepAL,一种面向完整视频选择的主动学习框架。StepAL融合了步骤感知特征表示,利用伪标签捕捉每段视频中预测步骤的分布,并结合熵加权聚类策略,优先选择既不确定又具备多样步骤组成的视频进行标注。在两个白内障手术数据集(Cataract-1k 和 Cataract-101)上的实验表明,StepAL始终优于现有主动学习方法,在更少标注视频条件下实现更高的步骤识别准确率。该方法为高效手术视频分析提供了有效路径,显著减轻了构建计算机辅助手术系统时的标注负担。
原文摘要 · Abstract (English)
Active learning (AL) can reduce annotation costs in surgical video analysis while maintaining model performance. However, traditional AL methods, developed for images or short video clips, are suboptimal for surgical step recognition due to inter-step dependencies within long, untrimmed surgical videos. These methods typically select individual frames or clips for labeling, which is ineffective for surgical videos where annotators require the context of the entire video for annotation. To address this, we propose StepAL, an active learning framework designed for full video selection in surgical step recognition. StepAL integrates a step-aware feature representation, which leverages pseudo-labels to capture the distribution of predicted steps within each video, with an entropy-weighted clustering strategy. This combination prioritizes videos that are both uncertain and exhibit diverse step compositions for annotation. Experiments on two cataract surgery datasets (Cataract-1k and Cataract-101) demonstrate that StepAL consistently outperforms existing active learning approaches, achieving higher accuracy in step recognition with fewer labeled videos. StepAL offers an effective approach for efficient surgical video analysis, reducing the annotation burden in developing computer-assisted surgical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。