研究乳腺钼靶中基础模型的偏见与泛化问题,发现数据不均衡导致性能差异。
Bias and Generalizability of Foundation Models across Datasets in Breast Mammography
- 对比不同数据集上预训练模型表现,揭示域间差异根源
- 聚合数据提升整体性能但未消除极端密度/年龄组偏差
- 公平性增强技术更稳定,适合医疗公平应用
过去几十年,乳腺癌辅助诊断工具旨在提升筛查效率,但其临床应用仍受数据变异性和固有偏见制约。尽管基础模型(FMs)通过利用大规模多样数据展现出强大泛化能力与迁移学习潜力,其性能仍可能因图像质量差异、标注不确定性及敏感患者特征引发虚假相关而受损。本文通过整合来自多元来源的大规模数据集(包括欠代表地区数据与自建数据集),系统评估乳腺钼靶分类中FMs的公平性与偏见。实验表明,模态特异性预训练可提升性能,但基于单一数据集特征训练的分类器在跨域时泛化失败;数据聚合虽改善整体表现,却未能完全缓解偏见,导致极重度乳腺密度与特定年龄群体出现显著性能差距。尽管领域自适应策略可减小差距,常伴随性能损失;相较之下,公平性感知方法在各子群体间表现更稳定且均衡。结果强调:构建包容且泛化的AI模型需融入严格的公平性评估与缓解机制。
原文摘要 · Abstract (English)
Over the past decades, computer-aided diagnosis tools for breast cancer have been developed to enhance screening procedures, yet their clinical adoption remains challenged by data variability and inherent biases. Although foundation models (FMs) have recently demonstrated impressive generalizability and transfer learning capabilities by leveraging vast and diverse datasets, their performance can be undermined by spurious correlations that arise from variations in image quality, labeling uncertainty, and sensitive patient attributes. In this work, we explore the fairness and bias of FMs for breast mammography classification by leveraging a large pool of datasets from diverse sources-including data from underrepresented regions and an in-house dataset. Our extensive experiments show that while modality-specific pre-training of FMs enhances performance, classifiers trained on features from individual datasets fail to generalize across domains. Aggregating datasets improves overall performance, yet does not fully mitigate biases, leading to significant disparities across under-represented subgroups such as extreme breast densities and age groups. Furthermore, while domain-adaptation strategies can reduce these disparities, they often incur a performance trade-off. In contrast, fairness-aware techniques yield more stable and equitable performance across subgroups. These findings underscore the necessity of incorporating rigorous fairness evaluations and mitigation strategies into FM-based models to foster inclusive and generalizable AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。