用生成模型扩充罕见皮肤病数据,提升诊断模型泛化能力。
One Shot GANs for Long Tail Problem in Skin Lesion Dataset using novel content space assessment metric

- 仅用一张样本生成新数据,解决罕见病数据不足问题。
- 在HAM10000数据集上显著提升少数类检测准确率。
- 设计专属评估指标,更真实反映生成数据质量。
医疗领域常面临长尾分布问题,尤其因罕见病症数据稀缺导致模型过拟合。当数据集类别极度不均衡时,模型易出现选择性检测——仅准确识别多数类而忽略少数类,从而降低泛化能力。为应对这一挑战,本文采用One Shot GANs对HAM10000数据集中尾部类别进行数据增强,生成额外样本以缓解不平衡问题。同时,提出一种专为One Shot GANs设计的新型内容空间评估指标,以更精准衡量生成数据质量与模型性能。实验表明,该方法有效提升了少数类的检测精度,增强了模型在新数据上的适用性。
原文摘要 · Abstract (English)
Long tail problems frequently arise in the medical field, particularly due to the scarcity of medical data for rare conditions. This scarcity often leads to models overfitting on such limited samples. Consequently, when training models on datasets with heavily skewed classes, where the number of samples varies significantly, a problem emerges. Training on such imbalanced datasets can result in selective detection, where a model accurately identifies images belonging to the majority classes but disregards those from minority classes. This causes the model to lack generalizability, preventing its use on newer data. This poses a significant challenge in developing image detection and diagnosis models for medical image datasets. To address this challenge, the One Shot GANs model was employed to augment the tail class of HAM10000 dataset by generating additional samples. Furthermore, to enhance accuracy, a novel metric tailored to suit One Shot GANs was utilized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。