用联邦学习在不共享数据前提下,提升多中心疼痛预测模型性能。
Privacy-preserving federated prediction of pain intensity change based on multi-center survey data
- 采用联邦学习在多地医疗数据间协作建模,数据始终留在本地。
- 丹麦数据中联邦模型R2达0.34,优于本地模型的0.30~0.32。
- 国际数据中联邦模型准确率0.78,显著高于本地模型的0.74。
患者自报调查数据用于训练预后模型以改善医疗。然而此类数据通常来自多中心,出于隐私考虑难以集中存储。本地训练模型准确性、鲁棒性和泛化能力较差。本文应用隐私保护的联邦机器学习技术构建预后模型,确保本地调查数据始终保留在医疗机构的安全范围内。我们在两个医疗数据集(丹麦五地区GLA:D数据和27国国际SHARE数据)上对比了集中式、本地和联邦学习方法,用于预测两种健康结局。比较了在线性回归、随机森林回归和随机森林分类模型上,本地训练与集中式及联邦训练的效果。在GLA:D数据中,联邦线性回归(R2 0.34,RMSE 18.2)和联邦随机森林回归(R2 0.34,RMSE 18.3)显著优于本地模型(R2 0.32,RMSE 18.6;R2 0.30,RMSE 18.8)。集中式模型(R2 0.34,RMSE 18.2;R2 0.32,RMSE 18.5)与联邦模型无显著差异。在SHARE数据中,联邦模型(准确率0.78,AUROC 0.71)和集中式模型(准确率0.84,AUROC 0.66)均显著优于本地模型(准确率0.74,AUROC 0.69)。结论:联邦学习可在不损害隐私的前提下,实现多中心调查数据的预后模型训练,性能损失极小甚至无损。
原文摘要 · Abstract (English)
Background: Patient-reported survey data are used to train prognostic models aimed at improving healthcare. However, such data are typically available multi-centric and, for privacy reasons, cannot easily be centralized in one data repository. Models trained locally are less accurate, robust, and generalizable. We present and apply privacy-preserving federated machine learning techniques for prognostic model building, where local survey data never leaves the legally safe harbors of the medical centers. Methods: We used centralized, local, and federated learning techniques on two healthcare datasets (GLA:D data from the five health regions of Denmark and international SHARE data of 27 countries) to predict two different health outcomes. We compared linear regression, random forest regression, and random forest classification models trained on local data with those trained on the entire data in a centralized and in a federated fashion. Results: In GLA:D data, federated linear regression (R2 0.34, RMSE 18.2) and federated random forest regression (R2 0.34, RMSE 18.3) models outperform their local counterparts (i.e., R2 0.32, RMSE 18.6, R2 0.30, RMSE 18.8) with statistical significance. We also found that centralized models (R2 0.34, RMSE 18.2, R2 0.32, RMSE 18.5, respectively) did not perform significantly better than the federated models. In SHARE, the federated model (AC 0.78, AUROC: 0.71) and centralized model (AC 0.84, AUROC: 0.66) perform significantly better than the local models (AC: 0.74, AUROC: 0.69). Conclusion: Federated learning enables the training of prognostic models from multi-center surveys without compromising privacy and with only minimal or no compromise regarding model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。