通过碰撞式水印实现更可靠的联邦学习后门检测
Coward: Collision-based OOD Watermarking for Practical Proactive Federated Backdoor Detection
- 利用多后门碰撞效应设计水印,主动干预训练过程
- 在多个基准数据集上达到当前最优检测性能
- 有效缓解分布外偏差问题,适合实际联邦学习场景
后门检测是当前防御联邦学习中后门攻击的主流方法,其中少数恶意客户端可上传中毒更新以破坏全局模型。现有方法分为被动和主动两类,但均存在实际局限:被动方法受非独立同分布数据分布和客户端随机参与影响,而现有主动方法因依赖后门共存效应,易受不可避免的分布外(OOD)偏差误导。为此,我们提出一种新主动检测方法Coward,基于多后门碰撞效应——连续植入的不同后门会显著抑制早期后门。相应地,通过在分布外数据上采用受控双映射学习,对联邦全局模型注入精心设计的后门碰撞水印。该设计不仅实现了与现有主动方法相反的检测范式,自然抵消了分布外预测偏差的负面影响,还引入低干扰训练干预,内在限制了分布外偏差强度,从而显著减少误判。大量实验表明,Coward在多个基准数据集上达到领先性能,并有效缓解分布外偏差。
原文摘要 · Abstract (English)
Backdoor detection is currently the mainstream defense against backdoor attacks in federated learning (FL), where a small number of malicious clients can upload poisoned updates to compromise the federated global model. Existing backdoor detection techniques fall into two categories, passive and proactive, depending on whether the server proactively intervenes in the training process. However, both of them have practical limitations: passive detection methods are disrupted by common non-i.i.d. data distributions and random participation of FL clients, whereas current proactive detection methods are misled by an inevitable out-of-distribution (OOD) bias because they rely on backdoor coexistence effects. To address these issues, we introduce a novel proactive detection method dubbed Coward, inspired by our discovery of multi-backdoor collision effects, in which consecutively planted, distinct backdoors significantly suppress earlier ones. Correspondingly, we modify the federated global model by injecting a carefully designed backdoor-collided watermark, implemented via regulated dual-mapping learning on OOD data. This design not only enables an inverted detection paradigm compared to existing proactive methods, thereby naturally counteracting the adverse impact of OOD prediction bias, but also introduces a low-disruptive training intervention that inherently limits the strength of OOD bias, leading to significantly fewer misjudgments. Extensive experiments on benchmark datasets show that Coward achieves state-of-the-art performance and effectively alleviates OOD bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。