arXiv:2509.20877cs.LG2025-09中稿 · the 2nd Workshop o…被引 1

通过控制客户端标签分布提升联邦学习性能

Distribution-Controlled Client Selection to Improve Federated Learning Strategies

  • 根据平衡或全局标签分布选择客户端
  • 本地数据不均衡时性能提升最明显
  • 适合数据分布差异大的联邦学习场景

联邦学习(FL)是一种分布式学习范式,允许多个客户端在保护数据隐私的前提下联合训练共享模型。尽管其在数据隐私要求严格的领域具有巨大潜力,但客户端间的数据不平衡会严重影响模型性能。为此,已有研究通过客户端选择方法缓解数据不平衡的负面影响。本文提出一种扩展的联邦学习策略,通过选择使当前标签分布最接近两种目标分布之一的活跃客户端:一是均衡分布,二是联邦整体合并后的标签分布。我们在三种常见联邦学习策略和两个数据集上进行了实证验证。结果表明,当面对局部不平衡时,与均衡分布对齐能带来最大性能提升;而在全局不平衡情况下,与联邦整体标签分布对齐表现更优。

原文摘要 · Abstract (English)

Federated learning (FL) is a distributed learning paradigm that allows multiple clients to jointly train a shared model while maintaining data privacy. Despite its great potential for domains with strict data privacy requirements, the presence of data imbalance among clients is a thread to the success of FL, as it causes the performance of the shared model to decrease. To address this, various studies have proposed enhancements to existing FL strategies, particularly through client selection methods that mitigate the detrimental effects of data imbalance. In this paper, we propose an extension to existing FL strategies, which selects active clients that best align the current label distribution with one of two target distributions, namely a balanced distribution or the federations combined label distribution. Subsequently, we empirically verify the improvements through our distribution-controlled client selection on three common FL strategies and two datasets. Our results show that while aligning the label distribution with a balanced distribution yields the greatest improvements facing local imbalance, alignment with the federation's combined label distribution is superior for global imbalance.

联邦学习客户端选择数据均衡隐私计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。