针对边缘物联网非独立同分布数据,提出多工作者选择算法提升分布式学习性能。
Multi-Worker Selection based Distributed Swarm Learning for Edge IoT with Non-i.i.d. Data
- 基于数据异质性度量,动态筛选对全局模型贡献大的边缘工作者
- 在多个异构数据集上验证,相比基准方法准确率提升5.2%~8.7%
- 适合边缘计算中数据分布不均的智能系统部署
分布式蜂群学习(DSL)为边缘物联网提供了增强数据隐私、通信效率与能效的前景。然而,非独立同分布(non-i.i.d.)数据会显著降低学习性能并引发训练发散。本文首次在DSL框架下量化非i.i.d.数据的影响,提出M-DSL算法,通过引入新的非i.i.d.度量指标,刻画本地数据集间的统计差异,并建立其与模型性能的关联。该机制指导高效选择对全局更新贡献突出的多个工作者。理论分析证明了算法收敛性,并在多种异构数据集和非i.i.d.设置下进行广泛实验。数值结果表明,相较于基准方法,本方案在准确率上提升5.2%~8.7%,显著增强网络智能。
原文摘要 · Abstract (English)
Recent advances in distributed swarm learning (DSL) offer a promising paradigm for edge Internet of Things. Such advancements enhance data privacy, communication efficiency, energy saving, and model scalability. However, the presence of non-independent and identically distributed (non-i.i.d.) data pose a significant challenge for multi-access edge computing, degrading learning performance and diverging training behavior of vanilla DSL. Further, there still lacks theoretical guidance on how data heterogeneity affects model training accuracy, which requires thorough investigation. To fill the gap, this paper first study the data heterogeneity by measuring the impact of non-i.i.d. datasets under the DSL framework. This then motivates a new multi-worker selection design for DSL, termed M-DSL algorithm, which works effectively with distributed heterogeneous data. A new non-i.i.d. degree metric is introduced and defined in this work to formulate the statistical difference among local datasets, which builds a connection between the measure of data heterogeneity and the evaluation of DSL performance. In this way, our M-DSL guides effective selection of multiple works who make prominent contributions for global model updates. We also provide theoretical analysis on the convergence behavior of our M-DSL, followed by extensive experiments on different heterogeneous datasets and non-i.i.d. data settings. Numerical results verify performance improvement and network intelligence enhancement provided by our M-DSL beyond the benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。