arXiv:2507.06844stat.MLcs.LG2025-07被引 1

动态选协作伙伴,让异构客户端更高效地个性化学习。

Adaptive collaboration for online personalized distributed learning with heterogeneous clients

  • 根据梯度相似性动态筛选协作同伴,降低更新方差。
  • 理论证明算法可加速收敛,速度取决于梯度方差大小。
  • 适合有异构数据的分布式个性化学习场景,如边缘计算。

我们研究了具有 N 个统计异构客户端的在线个性化去中心化学习问题,这些客户端通过协作加速本地训练。该设置中的关键挑战是选择相关协作者以减少梯度方差,同时控制引入的偏差。为此,我们提出一种基于梯度的协作准则,使每个客户端在优化过程中动态选择梯度相似的同伴。该准则源自对 All-for-one 算法的更精细、更通用的理论分析,其被 Even 等人(2022)证明为理想协作方案下的最优解。我们推导了平滑目标函数下的损失上界,适用于强凸、非凸或满足 Polyak-Lojasiewicz 条件的情况;分析表明,该算法起到方差缩减作用,加速效果依赖于足够的梯度方差。我们提出了两种实现该通用框架的协作方法,并证明其中一种变体保持了 All-for-one 的最优性。我们在合成和真实数据集上验证了结果。

原文摘要 · Abstract (English)

We study the problem of online personalized decentralized learning with $N$ statistically heterogeneous clients collaborating to accelerate local training. An important challenge in this setting is to select relevant collaborators to reduce gradient variance while mitigating the introduced bias. To tackle this, we introduce a gradient-based collaboration criterion, allowing each client to dynamically select peers with similar gradients during the optimization process. Our criterion is motivated by a refined and more general theoretical analysis of the All-for-one algorithm, proved to be optimal in Even et al. (2022) for an oracle collaboration scheme. We derive excess loss upper-bounds for smooth objective functions, being either strongly convex, non-convex, or satisfying the Polyak-Lojasiewicz condition; our analysis reveals that the algorithm acts as a variance reduction method where the speed-up depends on a sufficient variance. We put forward two collaboration methods instantiating the proposed general schema; and we show that one variant preserves the optimality of All-for-one. We validate our results with experiments on synthetic and real datasets.

个性化学习去中心化异构客户端梯度相似

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。