arXiv:2412.20341cs.LGcs.DC2024-12AAAI被引 20

解决联邦聚类中通信异步和聚类数未知的难题

Asynchronous Federated Clustering with Unknown Number of Clusters

  • 用种子点作为学习媒介,协调各客户端聚类
  • 自适应调整种子点,逐步揭示正确聚类数
  • 适合通信能力不均、聚类数未知的真实场景

联邦聚类(FC)在保护隐私的前提下,从多个客户端提供的非独立同分布(non-IID)无标签数据中挖掘知识至关重要。现有方法通常在本地学习聚类分布,再安全地将脱敏信息传至服务器聚合。然而,客户端通信能力差异以及聚类数 $k^*$ 未知等实际问题仍缺乏研究。本文首次揭示通信异步与未知 $k^*$ 存在复杂耦合关系,并提出异步联邦聚类学习(AFCL)方法。通过向客户端分发过多初始种子点作为学习媒介,跨客户端协同形成共识;为缓解因异步上传导致的分布不平衡,设计了种子更新平衡机制。实验表明,种子点可逐步自适应,准确揭示真实聚类数。

原文摘要 · Abstract (English)

Federated Clustering (FC) is crucial to mining knowledge from unlabeled non-Independent Identically Distributed (non-IID) data provided by multiple clients while preserving their privacy. Most existing attempts learn cluster distributions at local clients, and then securely pass the desensitized information to the server for aggregation. However, some tricky but common FC problems are still relatively unexplored, including the heterogeneity in terms of clients' communication capacity and the unknown number of proper clusters $k^*$. To further bridge the gap between FC and real application scenarios, this paper first shows that the clients' communication asynchrony and unknown $k^*$ are complex coupling problems, and then proposes an Asynchronous Federated Cluster Learning (AFCL) method accordingly. It spreads the excessive number of seed points to the clients as a learning medium and coordinates them across the clients to form a consensus. To alleviate the distribution imbalance cumulated due to the unforeseen asynchronous uploading from the heterogeneous clients, we also design a balancing mechanism for seeds updating. As a result, the seeds gradually adapt to each other to reveal a proper number of clusters. Extensive experiments demonstrate the efficacy of AFCL.

联邦学习聚类异步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。