通过分组分配带宽,加速工业物联网联邦学习训练
Bandwidth Allocation with Device Partitioning for Federated Learning over Industrial IoT networks

- 将设备分批有序分配全带宽,减少通信延迟
- 实测训练时间与能耗均低于现有方案,逼近理论最优
- 适合电池供电的边缘设备,提升工业场景实用性
我们研究工业物联网(IIoT)中联邦学习(FL)系统在无线信道上的协同训练问题,设备不共享本地数据,仅通过通信更新全局模型。由于通信耗时是主要瓶颈,传统按需分配带宽的方法难以满足高效收敛需求。本文提出一种基于设备计算能力异构性的新型带宽分配策略:将参与设备划分为有序子集,逐个子集独占全部带宽进行通信。理论上证明该策略在任何调度算法下均能实现比非分组方案更低的总训练时间。实验基于真实数据集(GC10-Det工业缺陷检测、CIFAR-10图像分类)验证,该方法显著降低训练时间和上行能耗,接近理论最小轮次时间,尤其适用于资源受限的电池驱动设备。
原文摘要 · Abstract (English)
We consider a federated learning (FL) system in which Industrial Internet-of-Things (IIoT) devices collaboratively train a global model over wireless channels without sharing local data. In such systems, communication time is a primary bottleneck that constrains overall training efficiency. Unlike conventional networks that prioritize individual quality-of-service requirements, FL systems collectively aim to converge to an optimal global model as efficiently as possible, which calls for a fundamentally different approach to bandwidth allocation. In this paper, we propose a novel bandwidth allocation policy that exploits the heterogeneity of device computing capabilities to minimize total training time. Rather than distributing bandwidth among all selected devices simultaneously, the proposed policy partitions the participating devices into ordered subsets and sequentially grants each subset exclusive access to the full bandwidth. We formally prove that this partitioning-based policy achieves a strictly lower training time than any bandwidth allocation scheme without partitioning, irrespective of the underlying scheduling algorithm. Furthermore, by reducing per-device transmission duration, the proposed policy also minimizes uplink energy consumption, which is particularly beneficial for battery-constrained IIoT devices. Extensive experiments on real-world datasets - including GC10-Det, an industrial surface defect benchmark, and CIFAR-10, a standard image classification benchmark - demonstrate that the proposed policy consistently reduces training time and energy consumption compared to existing bandwidth allocation schemes, approaching the theoretical lower bound on round time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。