用视觉变压器集成提升皮肤病变分类的可信度与跨域适应性
Domain Adaptive Skin Lesion Classification via Conformal Ensemble of Vision Transformers
- 构建视觉变压器集成框架,融合多数据集训练增强域适应能力
- 在跨域场景下实现90.38%覆盖率,较单模型提升9.95%
- 特别适合需要高置信度医疗决策支持的临床应用
探索深度学习模型的可信度在医疗影像决策支持系统等关键领域至关重要。共形预测为深度学习模型提供了可靠的不确定性估计和安全保证,但其性能受骨干模型在域偏移(如不同数据源差异)下表现不佳的影响。为此,本文提出一种新型框架——共形集成视觉变压器(CE-ViTs),通过优先考虑域适应性和模型鲁棒性来提升图像分类性能,并兼顾不确定性建模。该方法在HAME10000、Dermofit和Skin Cancer ISIC等多个数据集上训练视觉变压器集成,利用联合数据集校准,以增强共形学习中的域适应能力。实验表明,该框架实现了90.38%的高覆盖率,相比仅在HAM10000上训练的模型提升了9.95%,表明预测集包含真实标签的可能性显著提高。在难分类样本上,预测集平均大小从1.86增至3.075,显著提升共形预测性能。
原文摘要 · Abstract (English)
Exploring the trustworthiness of deep learning models is crucial, especially in critical domains such as medical imaging decision support systems. Conformal prediction has emerged as a rigorous means of providing deep learning models with reliable uncertainty estimates and safety guarantees. However, conformal prediction results face challenges due to the backbone model's struggles in domain-shifted scenarios, such as variations in different sources. To aim this challenge, this paper proposes a novel framework termed Conformal Ensemble of Vision Transformers (CE-ViTs) designed to enhance image classification performance by prioritizing domain adaptation and model robustness, while accounting for uncertainty. The proposed method leverages an ensemble of vision transformer models in the backbone, trained on diverse datasets including HAM10000, Dermofit, and Skin Cancer ISIC datasets. This ensemble learning approach, calibrated through the combined mentioned datasets, aims to enhance domain adaptation through conformal learning. Experimental results underscore that the framework achieves a high coverage rate of 90.38\%, representing an improvement of 9.95\% compared to the HAM10000 model. This indicates a strong likelihood that the prediction set includes the true label compared to singular models. Ensemble learning in CE-ViTs significantly improves conformal prediction performance, increasing the average prediction set size for challenging misclassified samples from 1.86 to 3.075.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。