arXiv:2605.09137cs.LG2026-05

FedAvg在乳腺密度异质性下仍保持高精度,适合医疗影像协同建模。

Evaluating Federated Learning approaches for mammography under breast density heterogeneity

论文配图:Evaluating Federated Learning approaches for mammography under breast density heterogeneity
图 1 · 摘自论文原文
  • 用联邦学习模拟不同密度分布的多中心数据,评估模型鲁棒性
  • FedAvg性能接近集中训练,优于本地模型和简单聚合方法
  • 无需特殊算法即可应对数据不平衡,适合临床实际部署

乳腺密度是影响乳腺钼靶解读的关键因素,也是多中心数据集中的主要异质性来源。这种异质性对跨机构协作机器学习(尤其是联邦学习)构成挑战。本研究评估了乳腺密度引起的异质性对乳腺钼靶图像分类中联邦学习的影响,并检验常见联邦学习算法在真实临床场景下的鲁棒性。实验设置两种情形:(1) 强异质性场景,各参与机构仅提供低密度(BI-RADS A-B)或高密度(BI-RADS C-D)病例;(2) 模拟白种人与亚洲人群乳腺密度分布的人群基准场景。在强异质性场景中,分别测试2客户端(分组为A-B vs C-D)和4客户端(每家仅一种密度)配置。对比三种联邦学习方法(FedAvg、FedProx、SCAFFOLD)与集中式训练、本地训练及朴素聚合方法(集成与加权平均)。在两种场景下,联邦学习性能均接近集中训练,而本地模型和朴素聚合方法在强异质性下表现较差。值得注意的是,FedAvg在准确率上达到甚至超过集中训练,证明其对乳腺密度相关数据不平衡具有天然鲁棒性,无需专门的异质性缓解算法。结果表明,联邦学习可有效应对乳腺密度异质性,支持其在真实乳腺钼靶工作流中的可行性。FedAvg的稳健表现凸显其在临床大规模部署的潜力,可在保障数据隐私的前提下实现协同建模。

原文摘要 · Abstract (English)

Breast density is a key factor that influences mammography interpretation and is a major source of heterogeneity in multicenter datasets. Such heterogeneity poses challenges for collaborative machine learning across institutions, particularly in Federated Learning. This study aims to evaluate the impact of breast density-induced heterogeneity on FL for mammography image classification and to assess the robustness of common FL algorithms in realistic clinical settings. We conducted experiments under two scenarios: (1) a strongly heterogeneous setting where each participating site contributed exclusively low- or high-density cases, based on the BI-RADS density score, and (2) a population-based setting simulating breast density distributions in White and Asian populations. For the strongly heterogeneous setting, we evaluated two configurations: one with 2 clients, where the cases were grouped as BI-RADS A-B and C-D, and one with 4 clients, where each site contained cases of a single BI-RADS density. We compared three FL methods (FedAvg, FedProx, SCAFFOLD) against centralized training, local-only training, and naive aggregation approaches, including ensembling and weight averaging. Across both scenarios, FL achieved performance comparable to centralized training, while local models and naive aggregation approaches underperformed in the presence of strong heterogeneity. Notably, FedAvg achieved accuracy on par with or exceeding centralized training, demonstrating resilience to breast density-induced data imbalance without requiring specialized heterogeneity mitigation algorithms. These findings show that FL can address breast density-related heterogeneity, supporting its feasibility for real-world mammography workflows. The demonstrated robustness of FedAvg underscores the potential for broad clinical deployment of FL, enabling collaborative model development while maintaining data privacy.

联邦学习乳腺癌筛查医学影像数据异质性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。