arXiv:2512.12737cs.LGcs.DC2025-12

提出SPARK方法,让非独立同分布的去中心化联邦学习更快更省通信。

Communication-Efficient Neural Tangent Kernels for Heterogeneous Decentralized Federated Learning

  • 用分阶段软标签正则化稳定动量加速神经切线核更新
  • 高异构下收敛快3倍,通信量降低70%达到目标精度
  • 适合数据异构严重且带宽受限的分布式场景

去中心化联邦学习(DFL)可在无中央服务器情况下协作训练模型,但在数据统计异构下收敛缓慢。已有研究表明,神经切线核(NTK)方法在DFL中比梯度更新收敛更快,而动量已被证明可加速梯度型联邦学习。然而,在异构数据下对NTK更新应用动量可能导致训练不稳定。本文提出SPARK,通过在邻域聚合数据上评估分阶段衰减的软标签正则化,使动量能稳定加速NTK更新。在高异构条件下,SPARK收敛速度比基线快约3倍,达到目标精度所需的总通信量最多降低70%,且在不同异构水平下均获得更高准确率。此外,我们研究了随机投影作为可选的雅可比矩阵压缩策略,适用于带宽受限场景。方法在多个数据集、网络拓扑和异构程度下均得到验证。

原文摘要 · Abstract (English)

Decentralized federated learning (DFL) enables collaborative model training without a central server, but converges slowly under statistical heterogeneity. Recent work has shown that neural tangent kernel (NTK) methods achieve faster convergence than gradient-based updates in DFL, while momentum has proven effective for accelerating gradient-based FL. However, applying momentum to NTK updates can destabilize training under heterogeneous data. We propose SPARK, which addresses this instability with a stage-wise annealed soft-label regularizer evaluated on neighborhood-aggregated data, so that momentum can accelerate NTK updates stably. Under high heterogeneity, SPARK converges about 3$\times$ faster than baselines and lowers the total communication to a target accuracy by up to about 70\%, and it attains higher accuracy across heterogeneity levels. We further study random projection as an optional Jacobian-compression strategy for bandwidth-constrained settings. We validate the approach across multiple datasets, network topologies, and heterogeneity levels.

联邦学习去中心化通信效率神经切线核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。