arXiv:2409.13004cs.LG2024-09被引 4

联邦学习中同时防范数据泄露与投毒攻击的动态扰动方法

Data Poisoning and Leakage Analysis in Federated Learning

  • 通过动态扰动模型实现隐私保护与抗投毒的协同防御
  • 实验证明该方法在保护隐私的同时维持模型性能
  • 适合关注联邦学习安全与隐私的开发者和研究者

数据投毒与泄露风险阻碍了联邦学习在现实中的大规模部署。本文揭示了两大主要威胁——训练数据隐私泄露与训练数据投毒——的真实情况与误区。首先分析了训练数据在联邦训练过程中可能泄露的时机与方式,提出一种在每轮联邦学习前对原始梯度更新添加可控随机噪声的防御策略。讨论了噪声量级与添加位置对缓解梯度泄露的关键作用。随后回顾并比较了多种数据投毒威胁,分析其引发模型后门攻击并损害全局模型性能的原因与条件。对代表性投毒攻击及其缓解技术进行分类对比,深入揭示数据投毒的负面影响。最后,展示动态模型扰动在同时保障隐私保护、抗投毒能力与模型性能方面的潜力。文章还讨论了联邦学习中的其他风险因素,如数据分布偏斜、数据与算法偏差及训练数据中的错误信息。基于实证证据,本研究为构建具备攻击抵御能力的联邦学习系统提供了变革性洞见。

原文摘要 · Abstract (English)

Data poisoning and leakage risks impede the massive deployment of federated learning in the real world. This chapter reveals the truths and pitfalls of understanding two dominating threats: {\em training data privacy intrusion} and {\em training data poisoning}. We first investigate training data privacy threat and present our observations on when and how training data may be leaked during the course of federated training. One promising defense strategy is to perturb the raw gradient update by adding some controlled randomized noise prior to sharing during each round of federated learning. We discuss the importance of determining the proper amount of randomized noise and the proper location to add such noise for effective mitigation of gradient leakage threats against training data privacy. Then we will review and compare different training data poisoning threats and analyze why and when such data poisoning induced model Trojan attacks may lead to detrimental damage on the performance of the global model. We will categorize and compare representative poisoning attacks and the effectiveness of their mitigation techniques, delivering an in-depth understanding of the negative impact of data poisoning. Finally, we demonstrate the potential of dynamic model perturbation in simultaneously ensuring privacy protection, poisoning resilience, and model performance. The chapter concludes with a discussion on additional risk factors in federated learning, including the negative impact of skewness, data and algorithmic biases, as well as misinformation in training data. Powered by empirical evidence, our analytical study offers some transformative insights into effective privacy protection and security assurance strategies in attack-resilient federated learning.

联邦学习隐私保护投毒攻击安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。