arXiv:2605.21563cs.LG2026-05

用嵌入式联邦学习预测缺铁,实现在不同医院的稳定部署。

Embedding-Based Federated Learning with Runtime Governance for Iron Deficiency Prediction

  • 用冻结的血液模型提取本地特征,只联邦训练小型分类器,减少通信开销。
  • 个性化聚合方法使两院区的预测准确率均提升,最大宏平均AUC达0.9133。
  • 支持医疗场景运行治理,实现权限控制与审计追踪,适合临床真实部署。

现有大多数医疗联邦学习研究难以落地。本文构建基于嵌入的联邦学习流程,利用常规全血细胞计数(FBC)数据预测缺铁,并在阿姆斯特丹大学医学中心(AUMC)与英国国家血液与移植机构(NHSBT)两个临床环境部署,二者缺铁患病率、铁蛋白分布及人群差异显著。采用冻结的领域专用血液基础模型DeepCBC进行本地表征提取,仅对下游小分类器进行联邦训练,大幅降低通信量。两个数据集非独立同分布(non-IID),差异源于人群异质性而非采样偏差。通过医疗导向的联邦平台FLA$^3$实现运行时治理,支持研究范围执行、策略化授权和签名审计日志。标准样本加权平均(FedAvg)因全局更新偏向更大的AUMC数据分布,导致两院区的受试者工作特征曲线下面积(ROC-AUC)低于本地训练。而个性化聚合方法FedMAP将AUMC的ROC-AUC从0.9470提升至0.9594,NHSBT从0.8558提升至0.8671,达到最高宏平均ROC-AUC 0.9133与最佳宏平均平衡准确率。结果表明,在客户端样本量与任务相关性差异显著的临床联邦中,个性化聚合更具优势。

原文摘要 · Abstract (English)

Recent reviews find that the vast majority of published healthcare federated learning (FL) studies never reach real-world deployment. We developed an embedding-based FL pipeline for iron deficiency prediction from routine full blood count (FBC) data and deployed it across real institutional environments at Amsterdam University Medical Centre (AUMC) and NHS Blood and Transplant (NHSBT), two clinical environments that differ markedly in iron deficiency prevalence, ferritin distribution, and subject populations. A frozen domain-specific haematology foundation model, DeepCBC, performs site-local representation extraction, restricting federated training to a compact downstream classifier and substantially reducing recurrent communication relative to full-encoder federation. The two clinical datasets are structurally not independent and identically distributed (non-IID), with heterogeneity arising from distinct population differences rather than sampling artefacts. Runtime governance is enforced by FLA$^3$, a healthcare-oriented FL platform providing study-scoped execution, policy-based authorisation, and signed audit logging. Standard sample-size-weighted aggregation (FedAvg) reduced the area under the receiver operating characteristic curve (ROC-AUC) at both sites relative to local-only training, as the global update was biased towards the larger AUMC distribution. FedMAP, a personalised aggregation method, raised ROC-AUC from 0.9470 to 0.9594 at AUMC and from 0.8558 to 0.8671 at NHSBT relative to local-only training, achieving the highest macro ROC-AUC of 0.9133 and the best macro balanced accuracy overall. These results support personalised aggregation in clinical federations where client sample size and task relevance diverge substantially.

联邦学习缺铁预测医疗AI个性化聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。