解决乳腺钼靶数据异质性问题,提升AI模型的可复现性与公平性。
MammoClean: Toward Reproducible and Bias-Aware AI in Mammography through Dataset Harmonization
- 统一数据筛选、图像处理和元数据格式,构建标准化框架。
- 在三个数据集上发现乳腺密度和病灶率存在显著分布差异。
- 证明数据污染会严重降低模型性能,适合医疗AI研究者使用。
开发临床可靠的乳腺钼靶人工智能系统受到公共数据集间数据质量、元数据标准和人群分布深度异质性的阻碍。这种异质性引入了数据集特异性偏差,严重损害模型泛化能力,成为临床部署的根本障碍。我们提出MammoClean,一个公开的乳腺钼靶数据集标准化与偏差量化框架。该框架统一病例筛选、图像处理(含侧别与亮度校正)及元数据结构。通过对乳腺解剖、成像特征及公共数据集的系统回顾,识别出关键偏差来源。将MammoClean应用于三个异构数据集(CBIS-DDSM、TOMPEI-CMMD、VinDr-Mammo),量化出乳腺密度和异常检出率的显著分布偏移。关键发现:训练数据受损的AI模型性能显著低于经清理后的模型。通过识别并缓解偏差源,研究人员可构建统一多数据集训练集,开发具备更强跨域泛化能力的鲁棒模型。MammoClean提供可复现的偏差感知AI开发流程,促进公平比较,推动安全、高效且在不同人群与临床场景中表现均衡的系统发展。开源代码已发布于:https://github.com/Minds-R-Lab/MammoClean。
原文摘要 · Abstract (English)
The development of clinically reliable artificial intelligence (AI) systems for mammography is hindered by profound heterogeneity in data quality, metadata standards, and population distributions across public datasets. This heterogeneity introduces dataset-specific biases that severely compromise the generalizability of the model, a fundamental barrier to clinical deployment. We present MammoClean, a public framework for standardization and bias quantification in mammography datasets. MammoClean standardizes case selection, image processing (including laterality and intensity correction), and unifies metadata into a consistent multi-view structure. We provide a comprehensive review of breast anatomy, imaging characteristics, and public mammography datasets to systematically identify key sources of bias. Applying MammoClean to three heterogeneous datasets (CBIS-DDSM, TOMPEI-CMMD, VinDr-Mammo), we quantify substantial distributional shifts in breast density and abnormality prevalence. Critically, we demonstrate the direct impact of data corruption: AI models trained on corrupted datasets exhibit significant performance degradation compared to their curated counterparts. By using MammoClean to identify and mitigate bias sources, researchers can construct unified multi-dataset training corpora that enable development of robust models with superior cross-domain generalization. MammoClean provides an essential, reproducible pipeline for bias-aware AI development in mammography, facilitating fairer comparisons and advancing the creation of safe, effective systems that perform equitably across diverse patient populations and clinical settings. The open-source code is publicly available from: https://github.com/Minds-R-Lab/MammoClean.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。