arXiv:2502.16545cs.CRcs.CV2025-02被引 1

通过特征聚合实现多目标联邦后门攻击,提升隐蔽性与成功率。

Multi-Target Federated Backdoor Attack Based on Feature Aggregation

  • 对齐触发器维度并限定像素边界,促进本地触发器间特征交互。
  • 零样本攻击成功率77.39%,可绕过现有防御检测。
  • 支持多目标同时生成触发器,降低攻击准备成本,适合安全评估者参考。

当前联邦后门攻击聚焦于协同训练后门触发器,多个受损客户端分别训练本地触发补丁,并在推理阶段合并为全局触发器。然而,这些方法需精心设计触发补丁的形状与位置,且训练中缺乏触发补丁间的特征交互,导致攻击成功率较低。此外,补丁像素未被裁剪,使后门样本在视觉上出现明显突变,易被检测算法发现。为此,我们提出基于特征聚合的新型联邦后门攻击基准。具体而言,我们将触发器维度与图像对齐,限定触发器像素边界,并促进各受损客户端训练的本地触发器之间的特征交互。同时,采用类内攻击策略,一次性生成所有目标类别的后门触发器,显著降低整体触发器生成时间,提高联邦模型被攻击的风险。实验表明,该方法不仅能绕过基于补丁的防御检测,还能实现77.39%成功率的零样本攻击。据我们所知,这是首个在联邦学习中实现此类零样本攻击的工作。最后,我们通过调整触发器训练因素(如投毒位置、比例、像素边界、本地训练轮数和通信轮数)评估了攻击性能。

原文摘要 · Abstract (English)

Current federated backdoor attacks focus on collaboratively training backdoor triggers, where multiple compromised clients train their local trigger patches and then merge them into a global trigger during the inference phase. However, these methods require careful design of the shape and position of trigger patches and lack the feature interactions between trigger patches during training, resulting in poor backdoor attack success rates. Moreover, the pixels of the patches remain untruncated, thereby making abrupt areas in backdoor examples easily detectable by the detection algorithm. To this end, we propose a novel benchmark for the federated backdoor attack based on feature aggregation. Specifically, we align the dimensions of triggers with images, delimit the trigger's pixel boundaries, and facilitate feature interaction among local triggers trained by each compromised client. Furthermore, leveraging the intra-class attack strategy, we propose the simultaneous generation of backdoor triggers for all target classes, significantly reducing the overall production time for triggers across all target classes and increasing the risk of the federated model being attacked. Experiments demonstrate that our method can not only bypass the detection of defense methods while patch-based methods fail, but also achieve a zero-shot backdoor attack with a success rate of 77.39%. To the best of our knowledge, our work is the first to implement such a zero-shot attack in federated learning. Finally, we evaluate attack performance by varying the trigger's training factors, including poison location, ratio, pixel bound, and trigger training duration (local epochs and communication rounds).

联邦学习后门攻击特征聚合零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。