实证研究联邦学习对模型准确率的影响,揭示其增益与退化场景。
An Empirical Study of the Impact of Federated Learning on Machine Learning Model Accuracy
- 在统一框架下系统测试文本、图像、音频、视频任务的联邦学习效果
- 发现部分场景下联邦学习使模型准确率大幅下降,部分场景影响可忽略
- 为实际部署和未来优化提供关键数据支持
联邦学习(FL)可在大规模私有用户数据上实现分布式机器学习模型训练。尽管已在多个领域展示潜力,但其对模型准确率的实际影响仍不清晰。本文通过系统性实证研究,考察这一学习范式对多种机器学习任务中前沿模型准确率的影响。实验涵盖文本、图像、音频、视频等多种数据类型,以及数据分布、联邦规模、客户端采样、本地与全局计算等配置参数。所有实验在统一的联邦学习框架中进行,投入大量人力与资源以确保高保真度。基于结果,我们进行了量化分析,识别出联邦学习导致模型准确率急剧下降的挑战性场景,以及影响可忽略的情况。详尽的发现有助于实际部署及未来联邦学习的发展。
原文摘要 · Abstract (English)
Federated Learning (FL) enables distributed ML model training on private user data at the global scale. Despite the potential of FL demonstrated in many domains, an in-depth view of its impact on model accuracy remains unclear. In this paper, we investigate, systematically, how this learning paradigm can affect the accuracy of state-of-the-art ML models for a variety of ML tasks. We present an empirical study that involves various data types: text, image, audio, and video, and FL configuration knobs: data distribution, FL scale, client sampling, and local and global computations. Our experiments are conducted in a unified FL framework to achieve high fidelity, with substantial human efforts and resource investments. Based on the results, we perform a quantitative analysis of the impact of FL, and highlight challenging scenarios where applying FL degrades the accuracy of the model drastically and identify cases where the impact is negligible. The detailed and extensive findings can benefit practical deployments and future development of FL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。