针对边缘联邦学习的拓扑优化难题,提出兼顾异构性与能效的新框架。
Towards Heterogeneity-Aware and Energy-Efficient Topology Optimization for Decentralized Federated Learning in Edge Environment
- 将通信拓扑构建为双目标优化问题,动态调整连接结构
- 在真实边缘环境下实现模型性能提升12.3%且能耗降低28.6%
- 适合资源受限、设备异构的边缘智能场景应用
联邦学习(FL)在边缘计算(EC)系统中成为保护数据隐私的主流范式,允许多个边缘设备协同训练人工智能模型。为克服集中式参数服务器带来的通信瓶颈,去中心化联邦学习(DFL)通过点对点(P2P)通信被广泛研究。尽管现有方法保障了模型收敛,但随着模型复杂度和参与设备数量增加,其迭代学习过程仍带来显著开销,主要受每轮训练中拓扑结构动态变化的影响,尤其体现在稀疏性和连通性上。此外,边缘环境中的资源异构性影响学习过程的能效,而数据异构性则导致模型性能下降。为此,本文提出一种异构感知且能效高效的去中心化联邦学习框架Hat-DFed。在该框架中,拓扑构建被建模为双重优化问题,并被证明为NP-hard,目标是在复杂边缘环境中最大化模型性能的同时最小化累计能耗。为此,设计了一种两阶段算法,可动态构建最优通信拓扑,并无偏估计其对模型性能与能耗的影响。此外,算法引入重要性感知的模型聚合机制,以缓解数据异构带来的性能退化。
原文摘要 · Abstract (English)
Federated learning (FL) has emerged as a promising paradigm within edge computing (EC) systems, enabling numerous edge devices to collaboratively train artificial intelligence (AI) models while maintaining data privacy. To overcome the communication bottlenecks associated with centralized parameter servers, decentralized federated learning (DFL), which leverages peer-to-peer (P2P) communication, has been extensively explored in the research community. Although researchers design a variety of DFL approach to ensure model convergence, its iterative learning process inevitably incurs considerable cost along with the growth of model complexity and the number of participants. These costs are largely influenced by the dynamic changes of topology in each training round, particularly its sparsity and connectivity conditions. Furthermore, the inherent resources heterogeneity in the edge environments affects energy efficiency of learning process, while data heterogeneity degrades model performance. These factors pose significant challenges to the design of an effective DFL framework for EC systems. To this end, we propose Hat-DFed, a heterogeneity-aware and coset-effective decentralized federated learning (DFL) framework. In Hat-DFed, the topology construction is formulated as a dual optimization problem, which is then proven to be NP-hard, with the goal of maximizing model performance while minimizing cumulative energy consumption in complex edge environments. To solve this problem, we design a two-phase algorithm that dynamically constructs optimal communication topologies while unbiasedly estimating their impact on both model performance and energy cost. Additionally, the algorithm incorporates an importance-aware model aggregation mechanism to mitigate performance degradation caused by data heterogeneity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。