arXiv:2412.05894q-bio.QMcs.CR2024-12

批量效应会干扰多中心组学联邦学习,需隐私保护校正。

Batch effects can impair federated learning in multi-center omics studies

  • 用联邦k均值与随机森林评估批量效应影响
  • 未校正批量效应使无监督学习失效,监督学习性能下降30%以上
  • 提出fedRBE实现隐私保护批量效应校正,支持缺失值和异构数据

联邦学习(FL)可在不共享患者级数据的前提下协作分析生物医学数据,但在多中心研究中可能受批量效应影响,掩盖真实生物信号。本文使用四个多中心组学数据集(转录组、蛋白质组、代谢组)及两种代表性算法(联邦k均值聚类和联邦随机森林分类),系统评估未校正批量效应的影响。结果表明,未校正的批量效应会破坏无监督联邦学习,显著降低监督学习性能。为在分布式群体组学数据中实现隐私保护的批量效应校正,我们提出fedRBE(https://featurecloud.ai/app/fedrbe),基于limma的removeBatchEffect()方法,结合安全多方计算,适用于含缺失值及特征集不一致的场景,包括蛋白质组与代谢组数据。

原文摘要 · Abstract (English)

Federated learning (FL) enables collaborative analysis of biomedical data without exchanging sensitive patient-level information, but its performance in multi-center studies may be compromised by batch effects which can obscure biological signals. Here, we systematically assess the impact of uncorrected batch effects on FL outcomes using four multi-center omics datasets, including transcriptomic, proteomic, and metabolomic data, and two representative algorithms: federated k-means clustering and federated random forest classification. Our results demonstrate that uncorrected batch effects undermine unsupervised FL and can substantially degrade supervised FL performance, indicating that privacy-aware batch-effect correction is essential for reliable FL. To enable privacy-preserving BEC in distributed bulk omics data, we introduce fedRBE ( https://featurecloud.ai/app/fedrbe ), a federated implementation of limma's removeBatchEffect() method enhanced by secure multi-party computation, suitable for datasets with missing values and non-identical feature sets across clients, including proteomics and metabolomics data.

联邦学习批量效应组学分析隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。