联邦学习下基于变分推断的贝叶斯聚类方法,保护隐私同时高效处理大规模数据。
Federated Variational Inference for Bayesian Mixture Models
- 分治式变分推断,局部并行处理数据块,全局合并聚类结构。
- 仅需各节点数据摘要即可完成联邦学习,无需共享原始数据。
- 适用于电子健康记录等敏感大数据的隐私保护聚类分析。
我们提出一种面向大规模二值与类别型数据的贝叶斯模型聚类联邦学习方法。引入基于变分推断的‘分而治之’推理流程,在数据批次内并行执行局部合并与删除操作,再通过跨批次的‘全局’合并操作识别整体聚类结构。这些合并操作仅需各批次的数据摘要信息,实现无需共享完整数据集的联邦学习。在模拟数据和基准数据集上的实验表明,该方法性能优于现有聚类算法。通过应用于大规模电子健康记录(EHR)数据,验证了其实际应用价值。
原文摘要 · Abstract (English)
We present a federated learning approach for Bayesian model-based clustering of large-scale binary and categorical datasets. We introduce a principled 'divide and conquer' inference procedure using variational inference with local merge and delete moves within batches of the data in parallel, followed by 'global' merge moves across batches to find global clustering structures. We show that these merge moves require only summaries of the data in each batch, enabling federated learning across local nodes without requiring the full dataset to be shared. Empirical results on simulated and benchmark datasets demonstrate that our method performs well in comparison to existing clustering algorithms. We validate the practical utility of the method by applying it to large scale electronic health record (EHR) data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。