arXiv:2409.18915cs.LG2024-09NeurIPS被引 6

解决联邦学习中长期不参与客户端的双变量漂移问题。

A-FedPD: Aligning Dual-Drift is All Federated Primal-Dual Learning Needs

  • 通过虚拟对偶更新对齐全局与本地对偶变量。
  • 在非凸目标下证明了优化与泛化效率优势。
  • 适合存在设备异构和部分参与的分布式场景。

作为平衡数据隐私与协同训练的主流范式,联邦学习(FL)正广泛应用于边缘客户端上分布处理大规模异构数据集。受带宽限制与安全考虑,其巧妙地将原问题拆分为多个子问题并行求解,使原始对偶方法在FL中具有重要应用价值。本文回顾经典联邦对偶方法的最新进展,指出其在非凸场景下的严重共性缺陷——由长期不活跃客户端引起的对偶滞后现象,即‘对偶漂移’。为解决该问题,我们提出新型对齐联邦对偶(A-FedPD)方法,通过构建虚拟对偶更新,对齐长期未参与客户端的全局共识与本地对偶变量。同时,我们对A-FedPD在光滑非凸目标下的优化与泛化效率进行了全面分析,证实其高效且实用。在多个经典联邦学习设置下进行了大量实验,验证了所提方法的有效性。

原文摘要 · Abstract (English)

As a popular paradigm for juggling data privacy and collaborative training, federated learning (FL) is flourishing to distributively process the large scale of heterogeneous datasets on edged clients. Due to bandwidth limitations and security considerations, it ingeniously splits the original problem into multiple subproblems to be solved in parallel, which empowers primal dual solutions to great application values in FL. In this paper, we review the recent development of classical federated primal dual methods and point out a serious common defect of such methods in non-convex scenarios, which we say is a "dual drift" caused by dual hysteresis of those longstanding inactive clients under partial participation training. To further address this problem, we propose a novel Aligned Federated Primal Dual (A-FedPD) method, which constructs virtual dual updates to align global consensus and local dual variables for those protracted unparticipated local clients. Meanwhile, we provide a comprehensive analysis of the optimization and generalization efficiency for the A-FedPD method on smooth non-convex objectives, which confirms its high efficiency and practicality. Extensive experiments are conducted on several classical FL setups to validate the effectiveness of our proposed method.

联邦学习对偶优化非凸优化分布式训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。