用扩散模型解耦数据分布,让联邦学习只需一次通信就高效运行
Disentangling data distribution for Federated Learning
- 用稳定扩散模型分离客户端数据分布,再重建
- 在CIFAR100和DomainNet上实现单轮通信下性能超越传统方法
- 适合关注隐私保护与通信效率的联邦学习研究者
联邦学习(FL)允许分布式客户端协作训练全局模型,利用私有数据提升性能的同时保障数据隐私。然而,客户端间数据分布的纠缠严重限制了其广泛应用。本文首次证明,通过解耦数据分布,联邦学习理论上可达到类似分布式系统的效率,仅需一轮通信。为此,我们提出新型FedDistr算法,采用稳定扩散模型对数据分布进行解耦与重建。在CIFAR100和DomainNet数据集上的实验表明,该方法在解耦及近解耦场景下均显著提升模型效用与效率,同时确保隐私安全,优于传统联邦学习方法。
原文摘要 · Abstract (English)
Federated Learning (FL) facilitates collaborative training of a global model whose performance is boosted by private data owned by distributed clients, without compromising data privacy. Yet the wide applicability of FL is hindered by entanglement of data distributions across different clients. This paper demonstrates for the first time that by disentangling data distributions FL can in principle achieve efficiencies comparable to those of distributed systems, requiring only one round of communication. To this end, we propose a novel FedDistr algorithm, which employs stable diffusion models to decouple and recover data distributions. Empirical results on the CIFAR100 and DomainNet datasets show that FedDistr significantly enhances model utility and efficiency in both disentangled and near-disentangled scenarios while ensuring privacy, outperforming traditional federated learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。