arXiv:2605.11122cs.CRcs.LG2026-05

通过关键层分析与代理替换,有效防御联邦学习中的后门攻击。

FedSurrogate: Backdoor Defense in Federated Learning via Layer Criticality and Surrogate Replacement

论文配图:FedSurrogate: Backdoor Defense in Federated Learning via Layer Criticality and Surrogate Replacement
图 1 · 摘自论文原文
  • 基于方向发散分析识别安全关键层,聚焦低维子空间检测异常。
  • 在非独立同分布数据下,误报率低于10%,攻击成功率控制在2.1%以下。
  • 用良性客户端的代理更新替换恶意梯度,保留梯度多样性。

联邦学习易受后门攻击影响——恶意客户端向全局模型注入特定行为。现有防御方法在真实非独立同分布(non-IID)数据下存在高误报率,错误标记良性客户端,导致模型准确率下降。我们提出FedSurrogate,结合双向梯度对齐过滤与层自适应异常检测,通过方向发散分析识别安全关键层,进行选择性聚类,将检测信号集中于低维子空间。双向软过滤阶段筛查可信客户端的残余污染,同时挽救被误判的良性客户端,显著降低异构条件下的误分类。不直接移除确认的恶意更新,而是用结构相似良性客户端的降维代理更新替代,保持梯度多样性并消除攻击影响。大量实验表明,FedSurrogate在所有数据集和攻击类型下,误报率均低于10%(相较最接近基线的31-32%显著降低),主任务准确率更优,且在挑战性non-IID设置下,攻击成功率始终低于2.1%。

原文摘要 · Abstract (English)

Federated Learning remains highly susceptible to backdoor attacks--malicious clients inject targeted behaviours into the global model. Existing defenses suffer from substantial false-positive rates under realistic non-independent and identically distributed (non-IID) data, incorrectly flagging benign clients and degrading model accuracy even when adversaries are correctly identified. We present FedSurrogate, a novel backdoor defense that addresses this limitation by combining bidirectional gradient alignment filtering with layer-adaptive anomaly detection. FedSurrogate performs selective clustering on security-critical layers identified via directional divergence analysis, concentrating the detection signal on a low-dimensional subspace. A bidirectional soft-filtering stage screens trusted clients for residual contamination while rescuing false positives from suspects, substantially reducing misclassifications under heterogeneous conditions. Rather than removing confirmed malicious updates, FedSurrogate replaces them with downscaled surrogate updates from structurally similar benign clients, preserving gradient diversity while neutralising adversarial influence. Extensive evaluations demonstrate that FedSurrogate maintains false-positive rates below 10% across all datasets and attack types, compared to 31-32% for the nearest comparably effective baseline, while achieving superior main-task accuracy and maintaining attack success rates below 2.1% across all tested datasets and attack types under challenging non-IID settings.

联邦学习后门攻击安全防御梯度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。