arXiv:2509.08089cs.LGcs.CR2025-09

提出防御联邦学习后门攻击的理论框架,区分大中小偏差攻击并组合防御方法。

Hammer and Anvil: Toward a Theory of Backdoors in Federated Learning

  • 按更新偏差δ分类后门攻击,设计大偏差检测与小偏差移除两类防御
  • 单一类型或非系统组合防御易被自适应攻击破解,单个恶意客户端即可成功
  • 联合使用两类防御可抵御最恶劣自适应攻击,实测在多种设置下完全有效

联邦学习(FL)支持分布式模型训练,但易受后门攻击:恶意客户端在全局模型中植入可控行为。现有防御对自适应攻击无效。本文提出“锤与砧”(Hammer and Anvil)理论框架,依据攻击更新与均值的偏差δ对后门进行分类。识别出两类根本性防御:类型1(砧)——基于异常值检测与鲁棒聚合,应对大偏差攻击;类型2(锤)——基于移除策略,应对小偏差攻击。我们证明,单一类型防御及非系统化组合防御均存在可被自适应攻击利用的漏洞。为此,提出类型1与类型2的系统性结合。在新设计的最坏情况、全信息自适应攻击者(知悉良性更新、聚合算法及参数)下,该组合防御仍失效。跨多个数据集和场景的实证表明,单一类型及非系统组合防御常被单个恶意客户端轻易攻破;而最优组合变体(HA_{Flame}^{CSFT}、HA_{Krum}^{CSFT}、HA_{Multi-Metrics}^{CSFT})在最极端对抗环境下仍保持安全。

原文摘要 · Abstract (English)

Federated Learning (FL) enables distributed model training but is vulnerable to backdoor attacks, where malicious clients embed attacker-controlled behaviors into the global model. Existing defenses fail against adaptive adversaries. In this paper, we present "Hammer and Anvil", a principled theoretical framework that categorizes backdoors by the deviation, $δ$, of their updates to the mean of the updates. We identify two fundamental defense types: "Type 1 (The Anvil)", comprising outlier detection and robust aggregation effective against large-deviation attacks, and "Type 2 (The Hammer)", consisting of removal-based defenses effective against small-deviation attacks. We demonstrate that defenses of a single type and non-principled combined defenses inherently leave an exploitable gap for adaptive attackers. To bridge this gap, we propose the principled combination of Type 1 and Type 2 defenses. We evaluate our framework against a new, worst-case, full-information adaptive adversary that knows the benign updates, the aggregation algorithm, and its parameters, and yet this adversary fails against our combined defenses. Our empirical evaluation across various datasets and settings shows that single-typed and non-principled combined defenses are easily broken, often by a single malicious client. In contrast, our best combined defense variants, $HA_{Flame}^{CSFT}$, $HA_{Krum}^{CSFT}$, and $HA_{Multi-Metrics}^{CSFT}$, remain undefeated even in the most adversarial settings. Our results provide a principled approach for research on backdoors in federated learning.

联邦学习后门攻击防御机制理论框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。