对比五种联邦学习方法在不均衡临床数据上的死亡预测表现
A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data
- 用联邦学习处理隐私敏感的异构临床数据,比较五种策略效果
- FedProx在AUC-ROC(0.897)和AUC-PR(0.230)上最佳,显著优于其他方法
- 不同科室模型表现差异大,需关注个体客户端性能而非仅整体指标
机器学习可预测院内死亡率,但数据隐私与临床数据统计异质性制约其应用。联邦学习(FL)具备隐私保护特性,但在非独立同分布(non-IID)和不平衡条件下表现仍需验证。本研究在MIMIC-IV数据集上,将466,351例住院记录按五个护理单元划分,构建真实非IID场景,并引入11项入院前24小时实验室指标作为特征。在1.98%的死亡率背景下,采用AUC-ROC与AUC-PR作为主要评价指标,避免使用阈值依赖的F1。经过50轮训练及五次随机种子实验,FedProx在AUC-ROC(0.897)和平均AUC-PR(0.230)上表现最优,配对t检验确认其显著领先于其他策略;但在F1指标上,FedCluster(0.280)略胜于FedProx(0.273),无单一策略全面占优。最优集中式基线模型的AUC-ROC和AUC-PR分别为0.929和0.312,显著优于FedProx。各客户端的分析显示,全局模型对不同护理单元的适配性不均(AUC-ROC 0.809–0.902),规模最小、临床最独特的客户端表现最差。结论:基于正则化的策略如FedProx更适用于联邦学习场景,而集中化仍具微弱预测优势,且应独立关注客户端异质性与评估集构建。
原文摘要 · Abstract (English)
Machine learning can predict in-hospital mortality, but data privacy and the statistical heterogeneity of clinical data hamper its use. Federated Learning (FL) is privacy-preserving, yet its behavior under non-IID and imbalanced conditions needs scrutiny. We benchmark five FL strategies - FedAvg, FedProx, FedAdagrad, FedAdam, and FedCluster - for mortality prediction on the MIMIC-IV dataset, partitioning 466,351 admissions across five care units to induce a realistic non-IID setting and enriching the features with an 11-item first-24-hour laboratory panel. At a prevalence of 1.98%, we adopt AUC-ROC and AUC-PR as primary, threshold-independent metrics rather than F1. Over 50 rounds and five random seeds, FedProx attains the best AUC-ROC (0.897) and mean AUC-PR (0.230), with paired t-tests confirming its AUC-ROC lead is significant against every other strategy; on F1, however, FedCluster (0.280) narrowly surpasses FedProx (0.273), so no single strategy dominates every metric. The best centralized baseline achieves AUC-ROC and AUC-PR of 0.929 and 0.312, respectively, significantly outperforming FedProx on both. A per-client breakdown shows the global model does not serve all care units equally (AUC-ROC 0.809-0.902), with the smallest, most clinically distinct clients faring worst. We conclude that regularization-based methods such as FedProx are the more robust federated choice, while centralization retains a slight predictive edge, and that per-client heterogeneity and evaluation-set construction deserve attention independent of aggregate numbers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。