精简医学影像数据集,提升对比学习的分类效果
Less is More: Selective Reduction of CT Data for Self-Supervised Pre-Training of Deep Learning Models with Contrastive Learning Improves Downstream Classification Performance
- 通过嵌入、信息论和哈希方法筛选并减少冗余数据
- 多个任务下游性能显著提升,如新冠CT分类AUC达0.83
- 加速预训练至原速的1/9,适合医疗图像模型优化
基于对比学习的自监督深度学习预训练在图像分析中广泛应用。现有研究显示其在医学图像领域具有巨大潜力,但需进一步结合医学图像特性。我们假设医学图像间相似性阻碍了对比学习的成功。为此,我们基于深度嵌入、信息论和哈希技术,探索不同策略以识别并减少医学预训练数据集中的冗余。评估了这些缩减策略在两个预训练数据集及多个下游分类任务上的表现。所有实验均显示,数据集缩减显著提升了下游任务性能:例如,新冠CT分类挑战赛的AUC从0.78提升至0.83,OrganSMNIST分类挑战赛从0.97升至0.98,脑出血分类任务从0.73增至0.83。同时,预训练速度最高提升九倍。结论表明,数据集质量至关重要,并为医学图像的对比学习预训练提供了可迁移的优化方法。
原文摘要 · Abstract (English)
Self-supervised pre-training of deep learning models with contrastive learning is a widely used technique in image analysis. Current findings indicate a strong potential for contrastive pre-training on medical images. However, further research is necessary to incorporate the particular characteristics of these images. We hypothesize that the similarity of medical images hinders the success of contrastive learning in the medical imaging domain. To this end, we investigate different strategies based on deep embedding, information theory, and hashing in order to identify and reduce redundancy in medical pre-training datasets. The effect of these different reduction strategies on contrastive learning is evaluated on two pre-training datasets and several downstream classification tasks. In all of our experiments, dataset reduction leads to a considerable performance gain in downstream tasks, e.g., an AUC score improvement from 0.78 to 0.83 for the COVID CT Classification Grand Challenge, 0.97 to 0.98 for the OrganSMNIST Classification Challenge and 0.73 to 0.83 for a brain hemorrhage classification task. Furthermore, pre-training is up to nine times faster due to the dataset reduction. In conclusion, the proposed approach highlights the importance of dataset quality and provides a transferable approach to improve contrastive pre-training for classification downstream tasks on medical images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。