用生成模型合成罕见肿瘤数据,缓解病理图像分类中的偏见问题。
MeDi: Metadata-Guided Diffusion Models for Mitigating Biases in Tumor Classification
- 将元数据融入扩散模型,生成针对少数群体的合成病理图像。
- 在TCGA数据上生成高质量新亚群图像,提升下游分类器性能。
- 适合关注医学图像公平性与数据增强的研究者使用。
近年来深度学习在组织学预测任务中取得显著进展,但其在临床应用中仍受限于对染色、扫描仪、医院及人口统计差异的鲁棒性不足:若训练数据偏向特定子群体,模型常出现捷径学习并产生偏见预测。大规模基础模型也未能完全解决此问题。为此,我们提出一种新型方法——元数据引导的生成式扩散模型(MeDi),通过将元数据显式建模,有针对性地为欠代表子群体生成合成数据,以平衡有限训练数据并缓解下游模型中的偏差。实验表明,MeDi可在TCGA中生成未见子群体的高质量病理图像,提升生成图像的整体保真度,并在存在子群体分布偏移的数据集上改善下游分类器性能。本工作为利用生成模型更好缓解数据偏见提供了概念验证。
原文摘要 · Abstract (English)
Deep learning models have made significant advances in histological prediction tasks in recent years. However, for adaptation in clinical practice, their lack of robustness to varying conditions such as staining, scanner, hospital, and demographics is still a limiting factor: if trained on overrepresented subpopulations, models regularly struggle with less frequent patterns, leading to shortcut learning and biased predictions. Large-scale foundation models have not fully eliminated this issue. Therefore, we propose a novel approach explicitly modeling such metadata into a Metadata-guided generative Diffusion model framework (MeDi). MeDi allows for a targeted augmentation of underrepresented subpopulations with synthetic data, which balances limited training data and mitigates biases in downstream models. We experimentally show that MeDi generates high-quality histopathology images for unseen subpopulations in TCGA, boosts the overall fidelity of the generated images, and enables improvements in performance for downstream classifiers on datasets with subpopulation shifts. Our work is a proof-of-concept towards better mitigating data biases with generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。