arXiv:2409.17189math.OCcs.LG2024-09被引 8

提出新算法让分布式学习在动态网络中高效收敛,支持自主调参。

Decentralized Federated Learning with Gradient Tracking over Time-Varying Directed Networks

  • 用梯度跟踪+动量机制实现跨节点协同优化
  • 精确梯度下线性收敛到最优解,随机梯度下趋近最优邻域
  • 支持异步步长与动量,适合实际部署中的灵活配置

研究在时变有向图上的去中心化(联邦)学习中的代理间交互问题,提出一种基于一致性思想的算法DSGTm-TV。该算法结合梯度跟踪与Heavy-Ball动量,通过行-列随机混合矩阵实现局部模型与梯度估计的分布式更新,在保障本地数据隐私的同时,确保共识与最优性。分析表明:当使用精确梯度时,DSGTm-TV可线性收敛至全局最优;采用随机梯度时,在期望意义下收敛至全局最优邻域。此外,相较于现有方法,该算法对非协调步长和动量参数仍保持收敛性,并给出显式边界。实验在真实世界图像分类与自然语言处理任务上验证了其有效性。

原文摘要 · Abstract (English)

We investigate the problem of agent-to-agent interaction in decentralized (federated) learning over time-varying directed graphs, and, in doing so, propose a consensus-based algorithm called DSGTm-TV. The proposed algorithm incorporates gradient tracking and heavy-ball momentum to distributively optimize a global objective function, while preserving local data privacy. Under DSGTm-TV, agents will update local model parameters and gradient estimates using information exchange with neighboring agents enabled through row- and column-stochastic mixing matrices, which we show guarantee both consensus and optimality. Our analysis establishes that DSGTm-TV exhibits linear convergence to the exact global optimum when exact gradient information is available, and converges in expectation to a neighborhood of the global optimum when employing stochastic gradients. Moreover, in contrast to existing methods, DSGTm-TV preserves convergence for networks with uncoordinated stepsizes and momentum parameters, for which we provide explicit bounds. These results enable agents to operate in a fully decentralized manner, independently optimizing their local hyper-parameters. We demonstrate the efficacy of our approach via comparisons with state-of-the-art baselines on real-world image classification and natural language processing tasks.

联邦学习去中心化梯度跟踪动态网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。