提出新方法防范联邦学习中的后门攻击,显著提升模型安全性。
Mitigating Backdoor Attacks in Federated Learning Using PPA and MiniMax Game Theory
- 结合声誉机制与博弈论动态识别恶意客户端
- 在多个数据集上将攻击成功率降至1.1%~11%
- 适合关注联邦学习安全性的研究人员和开发者
联邦学习因能利用分散数据并保护隐私而获得广泛应用,但其面临恶意客户端注入后门数据的问题,导致全局模型准确性和完整性受损。为应对这一威胁,本文提出FedBBA(联邦后门与行为分析)框架,通过声誉系统评估客户端行为、激励机制奖惩参与、结合投影追逐分析(PPA)与极小极大博弈理论,动态识别并抑制恶意客户端的影响。在德国交通标志识别基准(GTSRB)和比利时交通标志分类(BTSC)数据集上的大量仿真结果表明,该方法将后门攻击成功率降低至1.1%~11%,显著优于RDFL和RoPE等现有防御方案(攻击成功率23%~76%),同时保持正常任务准确率95%~98%。
原文摘要 · Abstract (English)
Federated Learning (FL) is witnessing wider adoption due to its ability to benefit from large amounts of scattered data while preserving privacy. However, despite its advantages, federated learning suffers from several setbacks that directly impact the accuracy, and the integrity of the global model it produces. One of these setbacks is the presence of malicious clients who actively try to harm the global model by injecting backdoor data into their local models while trying to evade detection. The objective of such clients is to trick the global model into making false predictions during inference, thereby compromising the integrity and trustworthiness of the global model on which honest stakeholders rely. To mitigate such mischievous behavior, we propose FedBBA (Federated Backdoor and Behavior Analysis). The proposed model aims to dampen the effect of such clients on the final accuracy, creating more resilient federated learning environments. We engineer our approach through the combination of (1) a reputation system to evaluate and track client behavior, (2) an incentive mechanism to reward honest participation and penalize malicious behavior, and (3) game theoretical models with projection pursuit analysis (PPA) to dynamically identify and minimize the impact of malicious clients on the global model. Extensive simulations on the German Traffic Sign Recognition Benchmark (GTSRB) and Belgium Traffic Sign Classification (BTSC) datasets demonstrate that FedBBA reduces the backdoor attack success rate to approximately 1.1%--11% across various attack scenarios, significantly outperforming state-of-the-art defenses like RDFL and RoPE, which yielded attack success rates between 23% and 76%, while maintaining high normal task accuracy (~95%--98%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。