用生成式AI增强数据,让异构物联网联邦学习更稳定高效
Generative AI-Powered Plugin for Robust Federated Learning in Heterogeneous IoT Networks
- 用生成式AI在边缘设备合成少数类数据,缓解数据分布不均
- 中心端筛选最接近均匀分布的设备参与训练,提升收敛速度与性能
- 适合数据稀疏或设备异构的隐私保护场景
联邦学习使边缘设备在本地保留数据的前提下协同训练全局模型,但设备间数据分布非独立同分布(Non-IID)常导致模型难收敛、性能下降。本文提出一种新型联邦优化插件,通过生成式AI增强的数据增强与均衡采样策略,将非独立同分布数据近似为独立同分布。核心思路是在各边缘设备上利用生成式AI合成欠代表类别数据,使整个联邦网络数据分布更均衡;同时,在中心服务器采用均衡采样机制,仅选择最接近独立同分布特性的设备参与训练,从而加速收敛并提升全局模型性能。实验验证该方法显著改善了收敛速度与抗数据不平衡能力,构建了一个灵活、隐私保护的联邦学习插件,适用于数据稀缺环境。
原文摘要 · Abstract (English)
Federated learning enables edge devices to collaboratively train a global model while maintaining data privacy by keeping data localized. However, the Non-IID nature of data distribution across devices often hinders model convergence and reduces performance. In this paper, we propose a novel plugin for federated optimization methods that approximates Non-IID data distributions to IID through generative AI-enhanced data augmentation and balanced sampling strategy. The key idea is to synthesize additional data for underrepresented classes on each edge device, leveraging generative AI to create a more balanced dataset across the FL network. Additionally, a balanced sampling approach at the central server selectively includes only the most IID-like devices, accelerating convergence while maximizing the global model's performance. Experimental results validate that our approach significantly improves convergence speed and robustness against data imbalance, establishing a flexible, privacy-preserving FL plugin that is applicable even in data-scarce environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。