arXiv:2412.03385cs.DCcs.LG2024-12ICML被引 2

动态调整分层联邦学习架构,平衡通信成本与模型精度。

Reactive Orchestration for Hierarchical Federated Learning Under a Communication Cost Budget

  • 基于多级监控实时感知变化,自动触发分层结构重构
  • 在通信预算内保持模型精度,动态优化资源利用效率
  • 支持云端边缘协同场景,适合大规模分布式训练

在计算连续体(CC)上部署分层联邦学习(HFL)需要将参与者组织成包含中间聚合节点的层级结构。但受通信成本约束、数据分布差异及运行环境波动等影响,实现难度大。为此,我们提出一种自适应编排框架,可对客户端流失和基础设施事件做出实时响应,同时权衡通信开销与模型精度。该框架利用多级监控信息(模型准确率、资源可用性、资源成本)识别需重构的事件,并引入通用的重构成本估算方法,持续评估适应动作质量,支持多种性能指标优化。通过扩展Kubernetes生态,系统能快速有效应对环境变化,在通信成本预算内实现运行时的成本与性能平衡。

原文摘要 · Abstract (English)

Deploying a Hierarchical Federated Learning (HFL) pipeline across the computing continuum (CC) requires careful organization of participants into a hierarchical structure with intermediate aggregation nodes between FL clients and the global FL server. This is challenging to achieve due to (i) cost constraints, (ii) varying data distributions, and (iii) the volatile operating environment of the CC. In response to these challenges, we present a framework for the adaptive orchestration of HFL pipelines, designed to be reactive to client churn and infrastructure-level events, while balancing communication cost and ML model accuracy. Our mechanisms identify and react to events that cause HFL reconfiguration actions at runtime, building on multi-level monitoring information (model accuracy, resource availability, resource cost). Moreover, our framework introduces a generic methodology for estimating reconfiguration costs to continuously re-evaluate the quality of adaptation actions, while being extensible to optimize for various HFL performance criteria. By extending the Kubernetes ecosystem, our framework demonstrates the ability to react promptly and effectively to changes in the operating environment, making the best of the available communication cost budget and effectively balancing costs and ML performance at runtime.

联邦学习边缘计算动态调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。