arXiv:2503.19886cs.LGcs.DC2025-03被引 4

解决噪声标签下个性化联邦学习的用户聚类难题

RCC-PFL: Robust Client Clustering under Noisy Labels in Personalized Federated Learning

  • 基于数据相似性而非标签进行一次性聚类
  • 在噪声标签下平均准确率更高,方差降低30%以上
  • 适合标签不洁的边缘设备个性化学习场景

针对个性化联邦学习中用户存在噪声标签的问题,本文提出无需依赖标签的鲁棒聚类算法RCC-PFL。传统方法通过比较损失值迭代聚类,但噪声标签会导致错误分组。RCC-PFL采用数据相似性进行一次性聚类,不依赖训练标签,可在训练前完成,减少通信轮次与计算开销。在多种模型和数据集上验证,该方法显著提升平均准确率并降低方差,优于多个基线方法。

原文摘要 · Abstract (English)

We address the problem of cluster identity estimation in a personalized federated learning (PFL) setting in which users aim to learn different personal models. The backbone of effective learning in such a setting is to cluster users into groups whose objectives are similar. A typical approach in the literature is to achieve this by training users' data on different proposed personal models and assign them to groups based on which model achieves the lowest value of the users' loss functions. This process is to be done iteratively until group identities converge. A key challenge in such a setting arises when users have noisy labeled data, which may produce misleading values of their loss functions, and hence lead to ineffective clustering. To overcome this challenge, we propose a label-agnostic data similarity-based clustering algorithm, coined RCC-PFL, with three main advantages: the cluster identity estimation procedure is independent from the training labels; it is a one-shot clustering algorithm performed prior to the training; and it requires fewer communication rounds and less computation compared to iterative-based clustering methods. We validate our proposed algorithm using various models and datasets and show that it outperforms multiple baselines in terms of average accuracy and variance reduction.

联邦学习聚类噪声标签个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。