通过模拟真实场景的图像增强,显著提升皮肤癌分类模型跨机构泛化能力。
Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification

- 设计基于物理成因的图像增强策略,减少不同设备/机构间数据差异影响。
- 在独立验证集上,模型ROC-AUC提升0.039,敏感度80%时特异度提高10个百分点。
- 增强策略不影响原有准确率,适合临床部署与多中心医疗系统应用。
皮肤病变分类模型在新诊所或新设备采集的数据上性能下降。本文研究何种数据增强方法可缓解这一问题,并采用策略选择与评估分离的测试协议。以六组皮肤镜数据(25,903张图像)训练一个ConvNeXt-Large二分类器,其中HAM10000和ISIC 2016-2020完全未参与训练。在1511张预留开发集上对单个增强、光度组合及11种复合策略进行排序;最优策略在8073张完全隔离的确认集上评估,该集合剔除了与训练数据共享病灶标识或来自训练机构的图像。每种策略用四个随机种子重训练,通过精确置换检验比较结果。结果显示,混合增强策略使确认集ROC-AUC从0.787升至0.826(+0.039),各种子区间不重叠(0.772–0.797 和 0.815–0.840),置换检验p=0.029。敏感度为0.80时特异度由0.612升至0.713,敏感度0.95时由0.284升至0.397。域内性能保持稳定(ROC-AUC 0.938→0.941)。在另一独立临床队列(472张图像,22例恶性)上表现也维持稳定(0.934 vs 0.930)。结论:模拟真实域偏移物理原因的增强策略能有效提升跨源迁移能力,且不影响域内性能,效果在无污染、分离评估下依然成立。
原文摘要 · Abstract (English)
Background/Objectives: Dermoscopic skin-lesion classifiers lose accuracy when images arrive from a new clinic or a new device. We asked which data augmentations reduce that loss, and measured the effect under a protocol that keeps policy selection separate from policy evaluation. Methods: A ConvNeXt-Large binary malignant-versus-non-malignant classifier was trained on six dermoscopic sources (25,903 images); HAM10000 and ISIC 2016-2020 were held out of training entirely. Single augmentations, photometric combinations and eleven composite policies were ranked on a development split of 1511 held-out images. The winning policy was then evaluated on a confirmation set of 8073 held-out images that took no part in that ranking and from which we removed every image sharing a lesion identifier with the training data and every image contributed by an institution represented in training. Both policies were retrained with four random seeds each and compared with an exact permutation test. Results: The mix policy raised confirmation-set ROC-AUC from 0.787 to 0.826 (+0.039; per-seed ranges 0.772-0.797 and 0.815-0.840, non-overlapping; exact permutation p=0.029), with the same direction on each contributing source. At matched sensitivity the gain is larger in clinical terms: specificity rose from 0.612 to 0.713 at a sensitivity of 0.80, and from 0.284 to 0.397 at a sensitivity of 0.95. In-domain ROC-AUC was preserved (0.938 to 0.941). On an independent clinical cohort acquired with a different device at a different institution (472 images, 22 malignant), performance was maintained (0.934 versus 0.930). Conclusions: Augmentations that model the physical causes of domain shift improve cross-source transfer at no cost to in-domain accuracy, and the improvement survives a selection-disjoint, contamination-free evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。