arXiv:2410.07678cs.LG2024-10被引 1

FedEP通过熵池化缓解去中心化联邦学习中的数据异构问题

FedEP: Tailoring Attention to Heterogeneous Data Distribution with Entropy Pooling for Decentralized Federated Learning

  • 用高斯混合模型拟合本地数据分布,共享统计参数估计全局分布
  • 基于局部与全局分布的熵池化计算聚合权重,提升收敛速度
  • 仅交换合成分布信息,兼顾隐私保护与通信效率,适合分布式场景

联邦学习中非独立同分布(non-IID)数据导致客户端漂移,影响收敛速度与模型性能。现有方法在集中式联邦学习(CFL)中有效缓解此问题,但去中心化联邦学习(DFL)因缺乏中心节点,难以获取全局视图,加剧了非IID挑战。本文受金融领域熵池化算法启发,提出联邦熵池化(FedEP)算法,利用高斯混合模型(GMM)拟合本地数据分布,通过邻居节点间共享统计参数估算全局分布。聚合权重由局部与全局分布间的熵池化方法确定。该方法仅传输合成分布信息,保障数据隐私并降低通信开销。实验表明,FedEP在多种非IID设置下均实现更快收敛,性能优于当前最优方法。

原文摘要 · Abstract (English)

Non-Independent and Identically Distributed (non-IID) data in Federated Learning (FL) causes client drift issues, leading to slower convergence and reduced model performance. While existing approaches mitigate this issue in Centralized FL (CFL) using a central server, Decentralized FL (DFL) remains underexplored. In DFL, the absence of a central entity results in nodes accessing a global view of the federation, further intensifying the challenges of non-IID data. Drawing on the entropy pooling algorithm employed in financial contexts to synthesize diverse investment opinions, this work proposes the Federated Entropy Pooling (FedEP) algorithm to mitigate the non-IID challenge in DFL. FedEP leverages Gaussian Mixture Models (GMM) to fit local data distributions, sharing statistical parameters among neighboring nodes to estimate the global distribution. Aggregation weights are determined using the entropy pooling approach between local and global distributions. By sharing only synthetic distribution information, FedEP preserves data privacy while minimizing communication overhead. Experimental results demonstrate that FedEP achieves faster convergence and outperforms state-of-the-art methods in various non-IID settings.

联邦学习去中心化数据异构熵池化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。