用熵引导的自监督学习提升医学图像分类精度
Entropy-Guided Self-Supervised Learning for Medical Image Classification
- 双模型协同:一个基于ImageNet预训练,一个用熵引导的MAE在医疗数据上预训练
- 在4个医学数据集上均达到顶尖水平,优于单模型和现有方法
- 适合需要高鲁棒性医学图像分类的研究者和临床应用
精准可靠的医学图像分类对早期疾病诊断和治疗规划至关重要。然而,标注数据有限、类内差异大、类间差异细微等问题常制约深度学习模型性能。本文提出一种融合自监督学习与迁移学习的协同框架:采用两个不同的ConvNeXt-Tiny模型——一个在大规模自然图像数据集ImageNet上预训练,另一个在目标医学数据集上通过熵引导的掩码自编码器(MAE)预训练——随后在特定医学图像分类任务上微调。最终采用概率平均的集成策略融合两模型互补信息。在四个多样化医学影像数据集(乳腺超声图像BUSI、ISIC 2018、Kvasir和COVID)上的严格实验验证表明,该集成方法性能卓越且鲁棒性强。基于MAE的预训练显著提升了领域特异性特征学习能力,而ImageNet预训练则提供强泛化特征。集成模型持续取得当前最优结果,证明结合多样预训练策略对挑战性医学图像分析的有效性。
原文摘要 · Abstract (English)
Accurate and robust medical image classification is paramount for early disease diagnosis and treatment planning. However, challenges such as limited annotated data, high intra-class variability, and subtle inter-class differences often hinder the performance of deep learning models. This paper introduces a synergistic deep learning framework that leverages the strengths of self-supervised learning and transfer learning for enhanced medical image classification. Our approach employs two distinct ConvNeXt-Tiny models: one pre-trained on a large-scale natural image dataset (ImageNet) and another pre-trained using an entropy-guided Masked Autoencoder (MAE) on the target medical dataset. Both models are then fine-tuned on specific medical image classification tasks. A final ensemble strategy, based on averaging predicted probabilities, is utilized to combine the complementary insights from these two models. Rigorous experimental validation across four diverse medical imaging datasets (Breast Ultrasound Images (BUSI), International Skin Imaging Collaboration (ISIC) 2018, Kvasir, and COVID) demonstrates the superior performance and robustness of our ensemble approach. The MAE pre-training significantly improves feature learning on domain-specific data, while the ImageNet pre-training provides strong generalizable features. The ensemble consistently achieves state-of-the-art results, outperforming individual models and existing methods, highlighting the efficacy of combining diverse pre-training strategies for challenging medical image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。