仅用视频级标签实现结肠镜视频中息肉定位,大幅降低标注成本。
WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos

- 利用视频级标签训练3D CNN生成激活图,定位息肉候选区域。
- 多视角策略提升小息肉定位准确率,[email protected]达30.97%。
- 适合医疗影像弱监督学习、内窥镜辅助诊断场景使用。
由于结肠镜视频的密集帧级标注成本高昂,本文提出WSPolypNet,一种仅需视频级标签的弱监督息肉定位框架。该框架采用3D卷积神经网络,在视频级监督下生成类别激活图(CAM),识别息肉候选区域,无需帧级空间标注。进一步通过多视角策略增强CAM定位线索,并作为点提示输入MedSAM2,由其在视频中传播分割掩码,根据息肉边界优化粗略定位。在IoU阈值为0.3、0.5、0.7时,CorLoc分别达到47.80%、43.68%、35.01%,较单视图设置提升显著;对小息肉,[email protected]从16.01%提升至30.97%。整体召回率达94.51%。结果表明,弱监督时空学习可显著减少结肠镜视频息肉定位的标注需求。
原文摘要 · Abstract (English)
Because dense frame-level annotation of colonoscopy videos is costly, we propose WSPolypNet, a weakly supervised framework for polyp localization using only video-level labels. WSPolypNet employs a 3D convolutional neural network trained with video-level supervision to generate class activation maps (CAMs), which identify candidate polyp regions without requiring frame-level spatial annotations. The CAM-derived localization cues are further enhanced using a multi-view strategy and provided to MedSAM2 as point prompts. MedSAM2 then propagates segmentation masks across the video, refining the coarse localization cues according to polyp boundaries. WSPolypNet achieved CorLoc scores of 47.80%, 43.68%, and 35.01% at IoU thresholds of 0.3, 0.5, and 0.7, respectively, compared with 36.87%, 33.72%, and 27.94% in the single-view setting. For small polyps, the multi-view strategy improved [email protected] from 16.01% to 30.97%. The framework also achieved a recall of 94.51%. These results demonstrate the potential of weakly supervised spatiotemporal learning to substantially reduce spatial annotation requirements for polyp localization in colonoscopy videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。