arXiv:2501.01913cs.CRcs.AI2025-01

提出新型后门攻击MIGO,让恶意模型更新伪装成正常数据,绕过防御持续植入后门。

Mingling with the Good to Backdoor Federated Learning

  • 设计攻击策略MIGO,使恶意更新自然融入合法更新中
  • 在5个数据集上后门准确率超90%,主任务性能不受损
  • 仅控制0.1%客户端仍可成功植入,适合研究防御漏洞者

联邦学习(FL)是一种分布式机器学习技术,允许多方在保护数据隐私的前提下联合训练模型。然而其分布式特性带来了诸多安全风险,已有多种防御机制通过多源数据与指标筛选恶意模型更新,以最小化或消除攻击影响。本文探讨了一种通用攻击方法的可行性,旨在实现后门植入的同时规避多样防御。聚焦于名为MIGO的攻击策略,该策略通过生成与合法更新高度相似的模型更新,实现后门的渐进式融合,使后门在攻击结束后仍能长期留存,并制造足够模糊性以干扰防御效果。MIGO在五个数据集和多种模型架构上植入三类后门,结果表明其后门准确率持续超过90%,且不影响主任务性能。同时,该攻击对十种防御手段表现出强规避能力,包括多项前沿方法。相较于四种其他攻击策略,MIGO在多数配置下均表现更优。值得注意的是,在极端场景下,即使攻击者仅控制0.1%的客户端,只要持续足够轮次,仍可成功植入后门。

原文摘要 · Abstract (English)

Federated learning (FL) is a decentralized machine learning technique that allows multiple entities to jointly train a model while preserving dataset privacy. However, its distributed nature has raised various security concerns, which have been addressed by increasingly sophisticated defenses. These protections utilize a range of data sources and metrics to, for example, filter out malicious model updates, ensuring that the impact of attacks is minimized or eliminated. This paper explores the feasibility of designing a generic attack method capable of installing backdoors in FL while evading a diverse array of defenses. Specifically, we focus on an attacker strategy called MIGO, which aims to produce model updates that subtly blend with legitimate ones. The resulting effect is a gradual integration of a backdoor into the global model, often ensuring its persistence long after the attack concludes, while generating enough ambiguity to hinder the effectiveness of defenses. MIGO was employed to implant three types of backdoors across five datasets and different model architectures. The results demonstrate the significant threat posed by these backdoors, as MIGO consistently achieved exceptionally high backdoor accuracy (exceeding 90%) while maintaining the utility of the main task. Moreover, MIGO exhibited strong evasion capabilities against ten defenses, including several state-of-the-art methods. When compared to four other attack strategies, MIGO consistently outperformed them across most configurations. Notably, even in extreme scenarios where the attacker controls just 0.1% of the clients, the results indicate that successful backdoor insertion is possible if the attacker can persist for a sufficient number of rounds.

联邦学习后门攻击模型安全欺骗性攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。