用SMOTETomek和FedProx提升不平衡医疗数据下差分隐私联邦学习的诊断效果
A Robust Pipeline for Differentially Private Federated Learning on Imbalanced Clinical Data using SMOTETomek and FedProx
- 在客户端使用SMOTETomek处理数据不平衡问题,提升少数类识别能力
- 优化后的FedProx在ε=9.0时仍保持77%以上召回率,显著优于标准FedAvg
- 为隐私与临床效用平衡提供可落地的技术路径,适合医疗联邦学习场景
联邦学习(FL)为协作医疗研究提供了新范式,可在保护患者隐私的同时实现分布式模型训练。当结合差分隐私(DP)时,虽能提供形式化安全保障,但常面临隐私与临床效用之间的权衡,尤其在医疗数据严重不平衡的场景下更为突出。本文针对心血管风险预测任务构建了多阶段分析框架:初始实验显示标准方法在不平衡数据上召回率为零。为此,我们在客户端集成混合过采样技术SMOTETomek,成功建立具有临床价值的模型;随后通过调优的FedProx算法应对非独立同分布(non-IID)数据挑战。最终结果揭示隐私预算(ε)与模型召回率之间存在非线性权衡,且优化后的FedProx始终优于标准FedAvg。在隐私-效用前沿上识别出最优操作区域:当ε=9.0时,仍可维持超过77%的召回率。本研究为真实世界异构医疗数据中高效、安全、精准的诊断工具开发提供了可复现的方法论蓝图。
原文摘要 · Abstract (English)
Federated Learning (FL) presents a groundbreaking approach for collaborative health research, allowing model training on decentralized data while safeguarding patient privacy. FL offers formal security guarantees when combined with Differential Privacy (DP). The integration of these technologies, however, introduces a significant trade-off between privacy and clinical utility, a challenge further complicated by the severe class imbalance often present in medical datasets. The research presented herein addresses these interconnected issues through a systematic, multi-stage analysis. An FL framework was implemented for cardiovascular risk prediction, where initial experiments showed that standard methods struggled with imbalanced data, resulting in a recall of zero. To overcome such a limitation, we first integrated the hybrid Synthetic Minority Over-sampling Technique with Tomek Links (SMOTETomek) at the client level, successfully developing a clinically useful model. Subsequently, the framework was optimized for non-IID data using a tuned FedProx algorithm. Our final results reveal a clear, non-linear trade-off between the privacy budget (epsilon) and model recall, with the optimized FedProx consistently out-performing standard FedAvg. An optimal operational region was identified on the privacy-utility frontier, where strong privacy guarantees (with epsilon 9.0) can be achieved while maintaining high clinical utility (recall greater than 77%). Ultimately, our study provides a practical methodological blueprint for creating effective, secure, and accurate diagnostic tools that can be applied to real-world, heterogeneous healthcare data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。