arXiv:2510.07132cs.LGcs.DC2025-10中稿 · ICASSP 2026被引 1

无需预设聚类数,自动发现最优客户端分组提升联邦学习性能

DPMM-CFL: Clustered Federated Learning via Dirichlet Process Mixture Model Nonparametric Clustering

  • 用狄利克雷过程先验实现无固定聚类数的非参数聚类
  • 在非独立同分布场景下,准确识别出隐藏的客户端分组结构
  • 适合数据异构性强、聚类未知的联邦学习应用场景

聚类联邦学习(CFL)通过将客户端分组并为每组训练专属模型,在非独立同分布(non-IID)环境下提升性能,平衡全局模型与完全个性化模型之间的权衡。然而,多数CFL方法需预先设定聚类数量K,当真实结构未知时极不实用。本文提出DPMM-CFL,采用狄利克雷过程(Dirichlet Process, DP)对聚类参数分布施加先验,实现非参数贝叶斯推断,可同时估计聚类数与客户端归属关系,并优化各组的联邦目标。该方法在每轮迭代中耦合联邦更新与聚类推断。实验在基准数据集上验证了其在狄利克雷分布和类别划分非独立同分布设置下的有效性。

原文摘要 · Abstract (English)

Clustered Federated Learning (CFL) improves performance under non-IID client heterogeneity by clustering clients and training one model per cluster, thereby balancing between a global model and fully personalized models. However, most CFL methods require the number of clusters K to be fixed a priori, which is impractical when the latent structure is unknown. We propose DPMM-CFL, a CFL algorithm that places a Dirichlet Process (DP) prior over the distribution of cluster parameters. This enables nonparametric Bayesian inference to jointly infer both the number of clusters and client assignments, while optimizing per-cluster federated objectives. This results in a method where, at each round, federated updates and cluster inferences are coupled, as presented in this paper. The algorithm is validated on benchmark datasets under Dirichlet and class-split non-IID partitions.

联邦学习聚类非参数贝叶斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。