研究联邦学习在持续数据流下的表现,证明其仍可有效协作学习。
Streaming Federated Learning with Markovian Data
- 采用小批量SGD、本地SGD及带动量的变体处理时间相关数据
- 样本复杂度与客户端数量成反比,通信开销接近独立同分布情况
- 适用于物联网、生物信号等实时数据场景
联邦学习(FL)是实现高效通信协同学习的关键框架。现有理论与实证研究多假设客户端拥有预先收集的数据集,对客户端持续采集数据的场景关注较少。在许多真实应用中,如物理或生物过程生成的数据,客户端数据流常被建模为非平稳马尔可夫过程。与标准独立同分布采样不同,由于样本间存在时间依赖性,马尔可夫数据流下的联邦学习性能尚不明确。本文研究联邦学习能否在马尔可夫数据流下仍支持协同学习。具体分析了小批量SGD、本地SGD及带动量的本地SGD变体。在标准假设和光滑非凸客户端目标下,答案为肯定:样本复杂度与客户端数量成反比,通信复杂度与独立同分布场景相当。但马尔可夫数据流的样本复杂度仍高于独立同分布采样。
原文摘要 · Abstract (English)
Federated learning (FL) is now recognized as a key framework for communication-efficient collaborative learning. Most theoretical and empirical studies, however, rely on the assumption that clients have access to pre-collected data sets, with limited investigation into scenarios where clients continuously collect data. In many real-world applications, particularly when data is generated by physical or biological processes, client data streams are often modeled by non-stationary Markov processes. Unlike standard i.i.d. sampling, the performance of FL with Markovian data streams remains poorly understood due to the statistical dependencies between client samples over time. In this paper, we investigate whether FL can still support collaborative learning with Markovian data streams. Specifically, we analyze the performance of Minibatch SGD, Local SGD, and a variant of Local SGD with momentum. We answer affirmatively under standard assumptions and smooth non-convex client objectives: the sample complexity is proportional to the inverse of the number of clients with a communication complexity comparable to the i.i.d. scenario. However, the sample complexity for Markovian data streams remains higher than for i.i.d. sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。