NebulaFL通过异步联邦学习提升多云协作效率,降低通信开销与成本。
NebulaFL: Effective Asynchronous Federated Learning for JointCloud Computing
- 采用版本控制的异步训练,缓解数据异构带来的性能下降。
- 通过去中心化模型轮转减少跨云通信开销,最高降50%。
- 结合奖励机制选源与资源调度,训练成本降低61.94%。
随着AI基础设施和可信执行环境(TEE)技术的发展,基于联合云计算(JCC)的联邦学习即服务(FLaaS)有望突破传统联邦学习中异构边缘设备带来的资源限制。在TEE保护下,数据所有者可利用云端高性能AI服务实现高效模型训练;云服务商通过提供额外的联邦学习服务,促进数据所有者间的协同学习。然而,FLaaS仍面临三大挑战:一、数据异构导致训练性能低下;二、不同云间通信开销高;三、缺乏有效资源调度策略以平衡训练时间和成本。为此,本文提出一种名为NebulaFL的新型异步联邦学习方法,用于多云间的协同模型训练。针对数据异构问题,NebulaFL在每个数据中心采用基于版本控制的异步训练方案,平衡各数据所有者的训练时长。为降低通信开销,引入去中心化模型轮转机制,实现数据中心间高效知识共享。为平衡训练时间与成本,集成基于奖励的参与者选择与资源调度策略。实验结果表明,相比现有最优方法,NebulaFL在目标准确率下,最高可提升5.71%准确率,通信开销降低50%,成本减少61.94%。
原文摘要 · Abstract (English)
With advancements in AI infrastructure and Trusted Execution Environment (TEE) technology, Federated Learning as a Service (FLaaS) through JointCloud Computing (JCC) is promising to break through the resource constraints caused by heterogeneous edge devices in the traditional Federated Learning (FL) paradigm. Specifically, with the protection from TEE, data owners can achieve efficient model training with high-performance AI services in the cloud. By providing additional FL services, cloud service providers can achieve collaborative learning among data owners. However, FLaaS still faces three challenges, i.e., i) low training performance caused by heterogeneous data among data owners, ii) high communication overhead among different clouds (i.e., data centers), and iii) lack of efficient resource scheduling strategies to balance training time and cost. To address these challenges, this paper presents a novel asynchronous FL approach named NebulaFL for collaborative model training among multiple clouds. To address data heterogeneity issues, NebulaFL adopts a version control-based asynchronous FL training scheme in each data center to balance training time among data owners. To reduce communication overhead, NebulaFL adopts a decentralized model rotation mechanism to achieve effective knowledge sharing among data centers. To balance training time and cost, NebulaFL integrates a reward-guided strategy for data owners selection and resource scheduling. The experimental results demonstrate that, compared to the state-of-the-art FL methods, NebulaFL can achieve up to 5.71\% accuracy improvement. In addition, NebulaFL can reduce up to 50% communication overhead and 61.94% costs under a target accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。