对比皮肤镜与临床图像数据差异,揭示影响模型性能的关键因素
An analysis of data variation and bias in image-based dermatological datasets for machine learning classification
- 分析皮肤镜与手机拍摄临床图像在分布上的差异
- 发现光照、肤色、视角等变化显著降低分类准确率
- 提出融合多源数据的策略以提升实际应用中的模型鲁棒性
AI算法在医疗领域日益重要,尤其在皮肤科中可基于RGB图像识别恶性病变。然而,多数模型依赖大型、标准的皮肤镜数据集训练,而真实临床场景使用智能手机拍摄,存在光照不稳、肤色差异、视角变化、噪声及标签不平衡等问题。尽管可采用迁移学习,但样本量少且训练与测试分布不一致导致性能下降。本文评估皮肤镜与临床图像间的分布差距,分析影响模型预测的关键差异,并通过多种架构实验,提出融合异构数据的策略,有效降低分布偏移对模型最终精度的影响。
原文摘要 · Abstract (English)
AI algorithms have become valuable in aiding professionals in healthcare. The increasing confidence obtained by these models is helpful in critical decision demands. In clinical dermatology, classification models can detect malignant lesions on patients' skin using only RGB images as input. However, most learning-based methods employ data acquired from dermoscopic datasets on training, which are large and validated by a gold standard. Clinical models aim to deal with classification on users' smartphone cameras that do not contain the corresponding resolution provided by dermoscopy. Also, clinical applications bring new challenges. It can contain captures from uncontrolled environments, skin tone variations, viewpoint changes, noises in data and labels, and unbalanced classes. A possible alternative would be to use transfer learning to deal with the clinical images. However, as the number of samples is low, it can cause degradations on the model's performance; the source distribution used in training differs from the test set. This work aims to evaluate the gap between dermoscopic and clinical samples and understand how the dataset variations impact training. It assesses the main differences between distributions that disturb the model's prediction. Finally, from experiments on different architectures, we argue how to combine the data from divergent distributions, decreasing the impact on the model's final accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。