arXiv:2409.03863cs.LG2024-09被引 3

首次理论量化本地更新对联邦学习泛化性能的影响。

Can We Theoretically Quantify the Impacts of Local Updates on the Generalization Performance of Federated Learning?

  • 基于线性模型分析,推导出泛化误差的闭式表达
  • 揭示本地更新次数K在三种情形下对泛化性能的影响
  • 为异构数据下的联邦学习提供新洞察,适合研究者参考

联邦学习(FL)因能有效在不共享数据的情况下跨多个站点训练模型而广受欢迎。尽管已有大量算法和优化分析表明,带有本地更新的联邦学习是通信高效的分布式学习框架,但其泛化性能却未得到充分关注。这主要归因于数据异构性与本地更新导致通信稀疏之间的复杂相互作用。本文提出一个根本性问题:能否在学习过程中量化数据异构性和本地更新对联邦学习泛化性能的影响?为此,我们以线性模型为起点,对联邦学习的泛化性能进行了全面的理论研究,涵盖平稳与在线/非平稳两种数据异构情形。通过给出模型误差的闭式表达,严格量化了本地更新次数K(K=1、K<∞、K=∞)的影响,并揭示了泛化性能随训练轮数t的变化规律。研究还系统分析了模型参数量p、样本数n等配置对整体泛化性能的贡献,为网络化联邦学习的实施提供了新见解(如良性过拟合现象)。

原文摘要 · Abstract (English)

Federated Learning (FL) has gained significant popularity due to its effectiveness in training machine learning models across diverse sites without requiring direct data sharing. While various algorithms along with their optimization analyses have shown that FL with local updates is a communication-efficient distributed learning framework, the generalization performance of FL with local updates has received comparatively less attention. This lack of investigation can be attributed to the complex interplay between data heterogeneity and infrequent communication due to the local updates within the FL framework. This motivates us to investigate a fundamental question in FL: Can we quantify the impact of data heterogeneity and local updates on the generalization performance for FL as the learning process evolves? To this end, we conduct a comprehensive theoretical study of FL's generalization performance using a linear model as the first step, where the data heterogeneity is considered for both the stationary and online/non-stationary cases. By providing closed-form expressions of the model error, we rigorously quantify the impact of the number of the local updates (denoted as $K$) under three settings ($K=1$, $K<\infty$, and $K=\infty$) and show how the generalization performance evolves with the number of rounds $t$. Our investigation also provides a comprehensive understanding of how different configurations (including the number of model parameters $p$ and the number of training samples $n$) contribute to the overall generalization performance, thus shedding new insights (such as benign overfitting) for implementing FL over networks.

联邦学习泛化性能理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。