在不共享数据的前提下,用联邦学习实现混合模型的自动聚类与参数推断。
Variational Consensus Monte Carlo for Bayesian Mixture

- 基于变分共识蒙特卡洛,实现各数据孤岛独立推断后聚合
- 可自动识别聚类数量,且在局部数据不完整时仍能准确匹配簇
- 适用于医疗等敏感数据场景,适合需保护隐私的研究者
针对医疗数据因隐私、敏感性及共享限制难以集中使用的问题,本文提出一种在联邦学习框架下推断贝叶斯混合模型的完整流程。采用共识蒙特卡洛(CMC)方法,在每个数据孤岛独立运行MCMC以估计局部后验分布,再通过聚合近似全局后验。现有变分CMC方法假设聚类数和关键参数已知,本文贡献包括:(i) 将变分CMC扩展至可自动推断聚类数和所有参数的过拟合贝叶斯混合模型,无需共轭假设;(ii) 提出适用于跨孤岛场景的新型聚类匹配算法,解决部分簇仅出现在局部数据的问题;(iii) 设计多种适配不同联邦约束的聚合推断策略;(iv) 提供实际选择指南。模拟实验验证了框架有效性,并与主流联邦学习方法对比。结果显示,当局部数据反映真实聚类结构时,本方法对小簇的识别精度优于将数据合并后使用标准MCMC。方法应用于英国老年群体大规模电子健康记录数据,成功识别多重共病模式。
原文摘要 · Abstract (English)
Motivated by the privacy, sensitivity and sharing limitations of health data, we present a comprehensive pipeline for inference of Bayesian mixture models within a federated learning setting, i.e. when data cannot be fully shared or pooled across compute nodes. We adopt a Consensus Monte Carlo (CMC) approach, in which an MCMC algorithm is run independently within each data silo to estimate local posterior distributions, which are then aggregated to approximate the posterior over the full data. The variational CMC approach of Rabinovich, Angelino and Jordan (2015) [1] frames the aggregation step as a variational inference problem, but their application to mixtures assumes the number of clusters and key mixture parameters to be known. Our main methodological contributions are: (i) an extension of variational CMC to over-fitted Bayesian mixture models that infer the number of clusters and all model parameters, without requiring conjugacy; (ii) novel cluster-matching algorithms suitable for cross-silo settings in which not every cluster appears in each local dataset; (iii) a number of inference strategies for the aggregation step, matched to different federated learning constraints; and (iv) guidelines for choosing among these in practice. A comprehensive simulation study validates the framework and allows us to compare to state-of-the-art federated learning alternatives. Notably, we show that when the composition of local datasets reflects the underlying clustering structure in the data, our approach can recover small clusters with greater accuracy than standard MCMC applied to the pooled data. We illustrate the framework on large-scale electronic health record data, identifying multi-morbidity patterns in a British geriatric population.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。