首个联邦学习生成时序电子病历的框架,解决医院数据隐私下联合建模难题。
FedEHR-Gen: Federated Synthetic Time-Series EHR Generation via Latent Space Alignment and Distribution-Aware Aggregation

- 分两阶段:先对齐各医院编码器到统一潜空间,再在潜空间上训练时间条件变分模型。
- 在eICU和MIMIC-III上生成质量接近集中式训练,且显著优于标准联邦基线。
- 适合需要跨院协作又严守数据隐私的医疗AI研发团队使用。
合成电子健康记录(EHR)为隐私受限的医疗场景下的数据增强与跨院建模提供了可行路径。然而,现有生成模型多为集中式,需汇聚各医院数据,实际难以实现。联邦学习虽可解决此问题,但因EHR数据高维、稀疏且跨院异质性强,直接联邦建模常出现崩溃或发散。本文提出首个跨分布式医院的时序EHR生成联邦框架FedEHR-Gen。该方法采用两阶段学习:首先设计联邦自编码器,将高维稀疏的EHR特征映射至紧凑潜空间;为保障语义一致性,提出逐层匹配聚合机制,对齐本地编码器至统一全局潜空间。其次,在对齐后的潜空间上训练联邦时间条件变分自编码器(TCVAE),并引入分布感知聚合策略,实现严重异质性下的稳定时序生成。在eICU和MIMIC-III数据集上的大量实验表明,FedEHR-Gen在生成保真度、下游任务效用及隐私风险方面均接近集中式训练表现,且持续优于标准联邦基线。
原文摘要 · Abstract (English)
Synthetic Electronic Health Record (EHR) generation provides a promising avenue for data augmentation and cross-hospital modeling in privacy-constrained healthcare settings. However, most existing EHR generative models are centralized and require pooling data across hospitals, which is often infeasible when real-world data sharing is restricted. While federated EHR generation offers a natural solution, direct federated modeling often collapses or diverges due to the high dimensionality, sparsity, and cross-hospital heterogeneity of EHR data. In this work, we propose FedEHR-Gen, the first federated framework for synthetic time-series EHR generation across distributed hospitals. FedEHR-Gen uses a two-stage learning paradigm. First, we introduce a federated autoencoder that projects high-dimensional and sparse EHR features onto a compact latent space. To ensure semantic consistency across hospitals, we develop a layer-wise matching aggregation mechanism that aligns local encoders into a unified global latent space. Second, operating on this aligned latent space, we train a federated temporal conditional variational autoencoder (TCVAE) with distribution-aware aggregation, enabling stable temporal generative modeling under severe cross-hospital heterogeneity. Extensive experiments on the eICU and MIMIC-III datasets demonstrate that FedEHR-Gen achieves generation fidelity, downstream utility, and privacy risk comparable to centralized training, while consistently outperforming the standard federated baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。