用多实例学习识别肠胶囊内镜中独特的息肉,提升诊断效率。
Using Multi-Instance Learning to Identify Unique Polyps in Colon Capsule Endoscopy Images
- 将息肉识别建模为多实例学习任务,通过注意力机制提取关键特征。
- 在1912个息肉数据上达到86.26%准确率和0.928的AUC。
- 结合自监督预训练,适合医学影像自动化分析研究者参考。
在肠胶囊内镜(CCE)图像中识别独特息肉是临床工作中的关键但具有挑战性的任务,因图像数量庞大、医生认知负荷高且标注存在模糊性。本文将该问题形式化为多实例学习(MIL)任务,以查询息肉图像与目标图像包对比判断其唯一性。采用融合注意力机制的多实例验证(MIV)框架,包括方差激发多头注意力(VEMA)和基于距离的注意力(DBA),增强特征表示能力。同时探索使用SimCLR进行自监督学习以生成鲁棒嵌入。在754名患者共1912个息肉的数据集上实验表明,注意力机制显著提升性能,采用ConvNeXt主干网络并经SimCLR预训练时,DBA L1达到最高测试准确率86.26%,测试AUC为0.928。研究证明了MIL与自监督学习在自动分析肠胶囊内镜图像中的潜力,对更广泛的医学影像应用具有启示意义。
原文摘要 · Abstract (English)
Identifying unique polyps in colon capsule endoscopy (CCE) images is a critical yet challenging task for medical personnel due to the large volume of images, the cognitive load it creates for clinicians, and the ambiguity in labeling specific frames. This paper formulates this problem as a multi-instance learning (MIL) task, where a query polyp image is compared with a target bag of images to determine uniqueness. We employ a multi-instance verification (MIV) framework that incorporates attention mechanisms, such as variance-excited multi-head attention (VEMA) and distance-based attention (DBA), to enhance the model's ability to extract meaningful representations. Additionally, we investigate the impact of self-supervised learning using SimCLR to generate robust embeddings. Experimental results on a dataset of 1912 polyps from 754 patients demonstrate that attention mechanisms significantly improve performance, with DBA L1 achieving the highest test accuracy of 86.26\% and a test AUC of 0.928 using a ConvNeXt backbone with SimCLR pretraining. This study underscores the potential of MIL and self-supervised learning in advancing automated analysis of Colon Capsule Endoscopy images, with implications for broader medical imaging applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。