用自监督预训练+领域适配,提升胃肠道内镜图像分类性能。
Domain-Adaptive Pre-training of Self-Supervised Foundation Models for Medical Image Classification in Gastrointestinal Endoscopy
- 基于EVA-02模型,在自建的EndoExtend24数据集上进行领域适配预训练。
- 在2024胶囊内镜挑战赛中取得宏平均AUC 0.762,平衡准确率37.1%。
- 数据集含超22万标注图像,支持123种病灶,适合多病种内镜诊断研究。
视频胶囊内镜已革新胃肠道内镜诊断,提供无创方式获取消化道高分辨率图像,实现疾病早期发现。但其应用受限于影像数量庞大(单次检查可生成多达100万张图像,持续6-8小时),需自动化分析。此外,图像差异大、专家标注成本高,且高质量标注数据集稀缺,制约现有医学图像分析模型效果。为此,本文构建了大型胃肠道内镜数据集EndoExtend24,整合十个公开与私有数据集,确保患者划分一致;该数据集包含超过226,000张标注图像,并支持动态类别映射,统一处理不同粒度标签,覆盖最多123种病理表现。我们提出利用自监督预训练的通用视觉基础模型(如基于ViT架构、在ImageNet-22k上通过掩码图像建模训练的EVA-02)进行领域适配预训练,以适应内镜诊断任务。具体地,将EVA-02在EndoExtend24上预训练后,再在胶囊内镜2024挑战赛数据集上微调。模型在挑战赛中获第三名,测试集上宏平均AUC达0.762,平衡准确率为37.1%,验证了该方法和数据集在推进胃肠道内镜诊断中的有效性。
原文摘要 · Abstract (English)
Video capsule endoscopy has transformed gastrointestinal endoscopy (GIE) diagnostics by offering a non-invasive method for capturing detailed images of the gastrointestinal tract, enabling early disease detection. However, its potential is limited by the sheer volume of images generated during the imaging procedure, which can take anywhere from 6-8 hours and often produce up to 1 million images, necessitating automated analysis. Additionally, the variability of these images, combined with the need for expert annotations and the scarcity of large, high-quality labeled datasets, constrains the effectiveness of current medical image analysis models. To address this, we introduce a novel large GIE dataset, called EndoExtend24, created by merging ten existing public and private datasets, ensuring patient integrity across splits. EndoExtend24 includes over 226,000 labeled images, as well as dynamic class mappings, which allow unified training across datasets with differing labeling granularity, supporting up to 123 distinct pathological findings. Further, we propose to leverage domain adaptive pre-training of foundation models trained with self-supervision on generic image data, to adapt them to the task of GIE medical image diagnosis. Specifically, the EVA-02 model, which is based on the ViT architecture and trained on ImageNet-22k with masked image modeling (using EVA-CLIP as a MIM teacher), is pre-trained on the EndoExtend24 dataset to achieve domain adaptation, and finally trained on the Capsule Endoscopy 2024 Challenge dataset. Our model demonstrates robust performance, securing third place in the Capsule Endoscopy 2024 Challenge. We achieved a macro AUC of 0.762 and a balanced accuracy of 37.1% on the test set. These results emphasize the effectiveness of our domain-adaptive pre-training approach and the enriched EndoExtend24 dataset in advancing gastrointestinal endoscopy diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。