解决去中心化学习中数据异构与隐私泄露问题
PDSL: Privacy-Preserved Decentralized Stochastic Learning with Heterogeneous Data Distribution
- 用谢林值衡量各节点对全局学习的贡献度
- 结合差分隐私保护梯度信息交换过程
- 适合注重隐私与数据异构场景的分布式学习
在去中心化学习中,多个智能体通过分布式数据协作训练全局模型,无需中心服务器。然而,各智能体间数据分布差异大,导致难以训练出鲁棒的全局模型;同时,智能体间依赖梯度信息交换,存在隐私泄露风险。本文提出PDSL算法,针对上述问题:一方面,引入谢林值思想,使每个智能体能精确评估其异构邻居对全局学习目标的贡献;另一方面,采用差分隐私机制,防止智能体在贡献梯度信息时遭受隐私泄露。我们进行了严谨的理论分析与广泛的实验,验证了PDSL在隐私保护与收敛性方面的有效性。
原文摘要 · Abstract (English)
In the paradigm of decentralized learning, a group of agents collaborates to learn a global model using distributed datasets without a central server. However, due to the heterogeneity of the local data across the different agents, learning a robust global model is rather challenging. Moreover, the collaboration of the agents relies on their gradient information exchange, which poses a risk of privacy leakage. In this paper, to address these issues, we propose PDSL, a novel privacy-preserved decentralized stochastic learning algorithm with heterogeneous data distribution. On one hand, we innovate in utilizing the notion of Shapley values such that each agent can precisely measure the contributions of its heterogeneous neighbors to the global learning goal; on the other hand, we leverage the notion of differential privacy to prevent each agent from suffering privacy leakage when it contributes gradient information to its neighbors. We conduct both solid theoretical analysis and extensive experiments to demonstrate the efficacy of our PDSL algorithm in terms of privacy preservation and convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。