跨机构联合训练心脏病风险模型,保护隐私且性能提升。
Federated Deep Learning for Privacy-Preserving Cardiovascular Disease Risk Prediction

- 采用联邦深度学习整合不同人群数据,不共享原始患者信息。
- 罗特丹研究C指数从0.728升至0.739,利夫林斯从0.783升至0.787。
- 适合需跨机构协作但受限于隐私法规的医疗预测任务。
心血管疾病风险预测模型通常依赖单一机构或集中数据集。跨机构扩展常受隐私法规和患者数据共享限制。联邦学习可在不传输敏感患者数据的前提下实现协作建模,但其在医疗领域的应用面临数据规模、人群特征和结局定义差异的挑战。本研究提出一种隐私保护的心血管疾病风险预测联邦深度学习方法,整合两个特征不同的队列:包含148,230名符合标准参与者的利夫林斯队列(自报结局)与包含10,155名参与者且结局由数字链接确认的罗特丹队列。模型性能主要在罗特丹队列上评估,因其随访完整。使用联邦学习训练的深度生存模型表现优于本地训练模型:罗特丹队列的C统计量从0.728(95% CI: 0.717–0.739)提升至0.739(95% CI: 0.728–0.749);利夫林斯队列从0.783(95% CI: 0.775–0.791)提升至0.787(95% CI: 0.780–0.792)。结果表明,跨异质队列的联邦深度学习可提升心血管疾病风险预测性能,同时保护个体数据隐私。
原文摘要 · Abstract (English)
Cardiovascular disease risk prediction models often rely on data from a single institution or centrally pooled datasets. Extending these models across institutions could be limited by privacy regulations and constraints on sharing patient-level data. Federated learning enables collaborative model development without transferring sensitive patient data, but its application in healthcare remains challenging because datasets often differ in size, population characteristics, and outcome definitions. In this study, we present a federated deep learning approach for privacy-preserving cardiovascular disease risk prediction that integrates two population-based cohorts with different characteristics: Lifelines, including 148,230 participants meeting the study inclusion criteria with self-reported outcomes, and the Rotterdam Study, including a smaller cohort of 10,155 participants with digitally linked clinical outcomes. Model performance was primarily evaluated on the Rotterdam Study because of its complete follow-up. Deep survival models trained using federated learning achieved higher predictive performance than models trained locally without federation. For the Rotterdam Study, the C-statistic increased from 0.728 (95% CI: 0.717-0.739) to 0.739 (95% CI: 0.728-0.749). For Lifelines, the C-statistic increased from 0.783 (95% CI: 0.775-0.791) to 0.787 (95% CI: 0.780-0.792). These findings suggest that federated deep learning across heterogeneous cohorts can improve cardiovascular disease risk prediction while preserving the privacy of individual-level patient data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。