跨机构聚类仅需一次通信,保护隐私还更准确。
Federated One-Shot Ensemble Clustering
- 各机构本地训练模型后,只传参数和标签,单轮通信完成聚类。
- 在类比例差异大的情况下仍保持聚类一致性,优于传统方法。
- 适合医疗等隐私敏感场景,特别适用于多中心研究。
跨机构聚类分析因数据共享限制面临重大挑战。为克服这些障碍,我们提出联邦一次性集成聚类(FONT)算法,专为多机构环境设计。该算法仅需一轮通信,通过交换拟合模型参数和类别标签实现隐私保护。其将本地拟合的聚类模型融合为数据自适应的集成模型,适用于多种聚类技术,且对各机构间簇比例差异具有鲁棒性。理论分析验证了FONT学习到的数据自适应权重的有效性,仿真研究显示其性能显著优于现有基准方法。我们将FONT应用于两个医疗系统中类风湿性关节炎患者的亚群识别,结果表明跨机构聚类一致性明显提升,而本地拟合聚类则转移性较差。该方法特别适用于通信与隐私要求严苛的真实场景,提供了一种可扩展、实用的多机构聚类解决方案。
原文摘要 · Abstract (English)
Cluster analysis across multiple institutions poses significant challenges due to data-sharing restrictions. To overcome these limitations, we introduce the Federated One-shot Ensemble Clustering (FONT) algorithm, a novel solution tailored for multi-site analyses under such constraints. FONT requires only a single round of communication between sites and ensures privacy by exchanging only fitted model parameters and class labels. The algorithm combines locally fitted clustering models into a data-adaptive ensemble, making it broadly applicable to various clustering techniques and robust to differences in cluster proportions across sites. Our theoretical analysis validates the effectiveness of the data-adaptive weights learned by FONT, and simulation studies demonstrate its superior performance compared to existing benchmark methods. We applied FONT to identify subgroups of patients with rheumatoid arthritis across two health systems, revealing improved consistency of patient clusters across sites, while locally fitted clusters proved less transferable. FONT is particularly well-suited for real-world applications with stringent communication and privacy constraints, offering a scalable and practical solution for multi-site clustering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。