arXiv:2409.14280math.OCcs.LG2024-09被引 1

融合海森与梯度相似性,降低分布式学习通信开销

Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning

  • 结合海森相似性与梯度一致性,设计新型加速算法
  • 实测在真实数据集上显著减少通信轮次
  • 适合大规模分布式训练与隐私敏感场景

现代学习任务对模型泛化能力要求越来越高,导致模型规模和训练样本量持续增长,单机训练已难胜任。因此分布式与联邦学习日益普及。这类方法涉及设备间通信,需解决效率与隐私两大问题。现有工作多单独研究局部数据相似性中的海森相似性或同质梯度,本文首次将两者结合,提出一种新方法,融合数据相似性与客户端采样思想,并引入额外噪声以保障隐私,分析其对收敛性的影响。理论分析得到真实数据集上的实验验证。

原文摘要 · Abstract (English)

Modern realities and trends in learning require more and more generalization ability of models, which leads to an increase in both models and training sample size. It is already difficult to solve such tasks in a single device mode. This is the reason why distributed and federated learning approaches are becoming more popular every day. Distributed computing involves communication between devices, which requires solving two key problems: efficiency and privacy. One of the most well-known approaches to combat communication costs is to exploit the similarity of local data. Both Hessian similarity and homogeneous gradients have been studied in the literature, but separately. In this paper, we combine both of these assumptions in analyzing a new method that incorporates the ideas of using data similarity and clients sampling. Moreover, to address privacy concerns, we apply the technique of additional noise and analyze its impact on the convergence of the proposed method. The theory is confirmed by training on real datasets.

分布式学习通信压缩联邦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。