arXiv:2501.03477cs.LG2025-01被引 1

分析联邦学习中通信瓶颈与数据非独立同分布问题对模型性能的影响

A study on performance limitations in Federated Learning

  • 在基础模型上分析通信开销和数据非独立同分布问题
  • 发现数据异构性显著降低模型准确率,通信效率制约训练速度
  • 适合关注隐私保护机器学习系统优化的研究者参考

日益增长的隐私担忧和数据无限制访问推动了联邦学习(Federated Learning, FL)这一新型机器学习范式的出现。FL借鉴分布式机器学习思想,但因其在边缘设备上训练模型而面临独特挑战。该范式由谷歌于2016年提出,此后在联邦优化算法、模型与更新压缩、差分隐私、鲁棒性与攻击防御、联邦GAN及隐私保护个性化等领域持续活跃研究。本文聚焦通信瓶颈与数据非独立同分布(Non-IID)对模型性能的影响,在基准模型上评估性能并探讨应对策略。

原文摘要 · Abstract (English)

Increasing privacy concerns and unrestricted access to data lead to the development of a novel machine learning paradigm called Federated Learning (FL). FL borrows many of the ideas from distributed machine learning, however, the challenges associated with federated learning makes it an interesting engineering problem since the models are trained on edge devices. It was introduced in 2016 by Google, and since then active research is being carried out in different areas within FL such as federated optimization algorithms, model and update compression, differential privacy, robustness, and attacks, federated GANs and privacy preserved personalization. There are many open challenges in the development of such federated machine learning systems and this project will be focusing on the communication bottleneck and data Non IID-ness, and its effect on the performance of the models. These issues are characterized on a baseline model, model performance is evaluated, and discussions are made to overcome these issues.

联邦学习隐私计算非独立同分布通信瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。