arXiv:2607.19524cs.LGcs.AI2026-07

用合成数据预训练联邦学习,提升医疗数据隐私下的预测鲁棒性。

SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework

论文配图:SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework
图 1 · 摘自论文原文
  • 先生成高保真合成电子病历,再用其预训练联邦模型。
  • 在5~15个异构客户端下,相比基线方法准确率显著提升。
  • 支持可解释性分析,适合医疗领域隐私敏感的分布式建模。

联邦学习(FL)为隐私保护的临床风险预测提供了前景,但受限于数据共享不足、客户端异构性、类别不平衡以及缺乏真实表格型电子健康记录(EHR)基准。合成数据生成可缓解数据稀缺,但其与联邦优化的结合尚未系统研究。本文提出SynPre-FL框架,将高保真合成EHR生成与合成预训练联邦学习相结合,以应对非独立同分布(non-IID)条件下的稳健预测。采用潜在空间自编码器-扩散模型生成隐私保护的合成队列,用于启动联邦训练。随后通过类别平衡本地目标、近端正则化及自适应服务器聚合实现异构性感知优化。事后校准与联邦安全可解释性支持可靠且可解释的风险估计。实验表明,合成生成器保持了单变量、双变量和多变量结构,有效抵御成员推断与重建攻击。生成数据在TSTR、TRTS及基于模型评估中表现出强下游效用。在5、10、15个异构客户端设置下,SynPre-FL持续优于基线方法,尤其在严重非IID条件下。校准提升了概率可靠性,SHAP分析在不同联邦规模下均产生稳定且符合临床逻辑的特征归因。SynPre-FL为此类场景提供了一种实用且可复现的隐私感知、可解释、鲁棒的临床预测框架。

原文摘要 · Abstract (English)

Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks. Synthetic data generation may alleviate data scarcity, yet its integration with federated optimisation has received limited systematic study. We propose SynPre-FL, a unified framework combining high-fidelity synthetic EHR generation with synthetic-pretrained FL for robust prediction under non-IID conditions. A latent autoencoder-diffusion model generates privacy-preserving synthetic cohorts, which are used to warm-start federated training. This pretraining is followed by heterogeneity-aware optimisation using class-balanced local objectives, proximal regularisation, and adaptive server aggregation. Post-hoc calibration and federated-safe explainability support reliable and interpretable risk estimates. Experiments show that the synthetic generator preserves univariate, bivariate, and multivariate structure while protecting against membership-inference and reconstruction attacks. The generated data achieve strong downstream utility under TSTR, TRTS, and model-based evaluations. Across federated settings with 5, 10, and 15 heterogeneous clients, SynPre-FL consistently improves robustness and scalability over baseline methods, especially under severe non-IID fragmentation. Calibration improves probability reliability, while SHAP analysis produces stable and clinically coherent feature attributions across federation sizes. SynPre-FL therefore provides a practical and reproducible framework for combining synthetic data with FL to enable privacy-aware, interpretable, and robust clinical prediction from distributed tabular EHR data.

联邦学习合成数据医疗预测隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。