FPLIER让多个机构在不共享数据的情况下联合训练基因通路分析模型。
FPLIER: Federated Pathway-Level Information Extractor

- 采用安全聚合技术实现跨机构分布式训练,数据本地化
- 通过增加数据秩使成员推断攻击失效,隐私风险显著降低
- 适合需要保护患者数据隐私的生物医学研究团队
在转录组学中,基于基因集的因子分解方法(如通路水平信息提取器,PLIER)在大规模异构表达数据集上表现最佳。然而,由于隐私和治理限制,许多临床相关队列无法合并为单一数据集。我们提出FPLIER,一种可扩展至多方协作的联邦版PLIER,支持在多个数据持有方之间分布式训练,同时融合公开数据集。通过安全聚合,FPLIER生成的训练更新在代数上等价于集中式合并数据的方法,且保持表达数据本地存储。我们在两个模拟联盟(来自K-CLIER和MultiPLIER研究)中评估了FPLIER,证明其具备稳定收敛性。进一步系统分析了针对中间训练统计和发布模型的成员推断攻击。结果表明,隐私风险由训练表达矩阵的秩决定:引入公开数据或降低数据维度可提升该秩,推动系统进入满秩状态,使得训练样本与非训练样本对攻击者不可区分,成员推断性能趋近随机猜测。
原文摘要 · Abstract (English)
In transcriptomics, gene-set-aware factorization methods such as the Pathway Level Information Extractor (PLIER) are most effective when trained on large, heterogeneous expression compendia. Yet, many clinically relevant cohorts cannot be pooled into a single dataset due to privacy and governance constraints. We present FPLIER, a federated extension of PLIER that enables distributed training across multiple data holders while incorporating publicly available datasets. Through secure aggregation, FPLIER produces training updates algebraically equivalent to those of a centralized pooled-data approach while keeping expression data local. We evaluate FPLIER across multiple scenarios in two simulated consortia (from the K-CLIER and MultiPLIER studies) and demonstrate stable convergence. We further conduct a systematic analysis of membership inference attacks targeting both intermediate training statistics and the released model. Our results show that privacy risk is governed by the rank of the training expression matrix. Incorporating public data or reducing data dimensionality increases this rank, moving the system toward a full-rank regime in which training and non-training samples become indistinguishable to the attacker, and membership-inference performance approaches random guessing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。