通过地理聚类分组,让联邦学习数据更均匀,提升模型性能。
Geographical Node Clustering and Grouping to Guarantee Data IIDness in Federated Learning
- 按设备地理位置动态聚类,分组时考虑移动性。
- 相比基准算法,联合成本降低至少110倍,分组数仅增0.93个。
- 适合大规模移动物联网场景下的联邦学习应用。
联邦学习(FL)是一种适用于智能物联网中大量设备的去中心化人工智能机制。其主要挑战是数据非独立同分布(non-IID)问题,源于参与方采集数据的异质性,导致全局模型性能下降。现有方法多聚焦于数据操作以缓解非IID问题,本文提出新思路:利用设备地理特性对移动物联网节点进行合理聚类与分组,确保每组数据接近独立同分布(IID)。首先实证验证了设备间距离与数据独立同分布特征的相关性;随后提出动态聚类与部分稳定分组算法,在考虑设备移动性的前提下,实现各组数据近似IID。该机制在联合成本(包括掉线设备数与组内设备数量均衡性)上显著优于基准分组算法,最低提升达110倍,分组数量最多仅增加0.93组。
原文摘要 · Abstract (English)
Federated learning (FL) is a decentralized AI mechanism suitable for a large number of devices like in smart IoT. A major challenge of FL is the non-IID dataset problem, originating from the heterogeneous data collected by FL participants, leading to performance deterioration of the trained global model. There have been various attempts to rectify non-IID dataset, mostly focusing on manipulating the collected data. This paper, however, proposes a novel approach to ensure data IIDness by properly clustering and grouping mobile IoT nodes exploiting their geographical characteristics, so that each FL group can achieve IID dataset. We first provide an experimental evidence for the independence and identicalness features of IoT data according to the inter-device distance, and then propose Dynamic Clustering and Partial-Steady Grouping algorithms that partition FL participants to achieve near-IIDness in their dataset while considering device mobility. Our mechanism significantly outperforms benchmark grouping algorithms at least by 110 times in terms of the joint cost between the number of dropout devices and the evenness in per-group device count, with a mild increase in the number of groups only by up to 0.93 groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。