研究去中心化梯度下降在马尔可夫采样下的泛化能力
Stability and Generalization for Decentralized Markov SGD

- 基于稳定性框架分析马尔可夫依赖与分布式通信的联合影响
- 给出去中心化SGD和SGDA的非渐近泛化界,首次覆盖最小最大场景
- 适用于分布式强化学习、联邦学习等存在数据相关性的场景
随机梯度方法是大规模学习的核心,但其泛化理论通常依赖独立采样假设。在许多实际应用中,数据由马尔可夫链生成,且学习过程为去中心化模式,这带来了显著的分析挑战。本文研究了在马尔可夫链采样下,去中心化随机梯度下降(SGD)和随机梯度下降上升(SGDA)的稳定性和泛化性。通过稳定性框架,我们刻画了马尔可夫依赖性与去中心化通信共同对泛化行为的影响。分析涵盖了网络拓扑、马尔可夫链混合性质以及原始-对偶动态。建立了两类算法的非渐近泛化边界,将现有马尔可夫随机梯度方法的结果拓展至去中心化及极小极大设置。
原文摘要 · Abstract (English)
Stochastic gradient methods are central to large-scale learning, yet their generalization theory typically relies on independent sampling assumptions. In many practical applications, data are generated by Markov chains and learning is performed in a decentralized manner, which introduces significant analytical challenges. In this work, we investigate the stability and generalization of decentralized stochastic gradient descent (SGD) and stochastic gradient descent ascent (SGDA) under Markov chain sampling. Leveraging a stability-based framework, we characterize how Markovian dependence and decentralized communication jointly influence generalization behavior. Our analysis captures the effects of network topology, Markov chain mixing properties, and primal-dual dynamics. We establish non-asymptotic generalization bounds for both algorithms, extending existing results on Markov stochastic gradient methods to decentralized and minimax settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。