评估不完美数据对联邦学习客户端选择的影响并提出隐私保护评分方法
Assessing the Impacts of Imperfect Datasets on Client Selections in Federated Learning

- 通过实验量化非独立同分布、噪声数据及选择偏差对模型性能的影响
- 提出隐私保护评分机制,有效评估客户端贡献度并提升整体模型精度
- 适合关注联邦学习中数据质量与公平性问题的研究者参考
联邦学习(FL)是一种分布式学习框架,多个客户端进行本地训练,服务器聚合本地更新的模型,可在保护客户端数据隐私的同时实现去中心化训练。然而,非独立同分布(non-IID)或含噪声的数据可能导致模型准确率下降或收敛延迟。通过客户端选择排除此类客户端可缓解问题,但过度偏倚的选择也可能损害学习性能。本研究首先实验评估了数据量与标签分布偏移、噪声数据以及客户端选择公平性对模型准确率和收敛速度的影响。随后提出一种隐私保护的客户端贡献评分方法,并通过实验验证其有效性。
原文摘要 · Abstract (English)
Federated learning (FL) is a popular distributed learning framework where multiple clients perform local training and a server aggregates the locally updated models. FL enables decentralized training while preserving the privacy of clients' datasets. However, non-independent and identically distributed (non-IID) or noisy datasets can lead to low model accuracy or high convergence latency. Precluding these clients through client selection may mitigate the problem, but heavily biased client selections may also degrade the learning performance. In this study, we first experimentally measure the impact of non-IID data (including skews in data quantity and label distribution), noisy data, and fairness in client selection on model accuracy and convergence. We then propose a privacy-preserving scoring method to assess each client's contribution in FL, with experiments conducted to demonstrate the effectiveness of the proposed assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。