arXiv:2508.04064cs.LGcs.AI2025-08被引 1

提出新方法检测联邦学习中隐藏的后门攻击,发现单一触发器测试会漏掉多触发模式。

FLAT: Revealing Hidden Latent-Conditioned Backdoor Failures in Federated Learning

  • 设计隐变量分离的应力测试,揭示多触发器激活同一目标的隐蔽攻击
  • 在多个数据集上保持高清洁准确率的同时,单目标攻击成功率超94%
  • 适合关注联邦学习安全性的研究人员,尤其关注防御盲区

水平联邦学习(HFL)的后门审计通常通过干净准确率(CA)、平均攻击成功率(ASR)或单一已知触发器测试来总结模型行为。此类总结可能掩盖一种新型失败模式:一个目标标签可被多种触发器实现激活。本文提出FLAT,一种针对HFL后门的隐变量条件可靠性压力测试。在该测试中,受控客户端仍提交常规分类器更新,而攻击端生成器 $G(x,t,z)$ 将目标意图 $t$ 与触发器实现 $z$ 分离。这一分离将审计问题从‘单一已知触发器是否成功’转变为‘隐藏行为如何随目标、隐变量样本、防御策略和停止后轮次变化’。在CIFAR-10、CIFAR-100和Tiny-ImageNet上,FLAT在保持干净效用的同时,实现了99.49%、99.66%和94.10%的单目标FedAvg ASR。评估还发现防御响应非均匀:服务器规则可抑制某一目标模式,却让另一模式持续活跃。这些观察促使我们采用更细致的审计指标,如目标级ASR、最差目标ASR、目标覆盖率、隐变量采样行为、停止后持久性及防御响应。

原文摘要 · Abstract (English)

Horizontal federated learning (HFL) backdoor audits often summarize model behavior through clean accuracy (CA), mean attack success rate (ASR), or a single known-trigger test. Such summaries can hide a different failure mode, in which one target label is activated by many trigger realizations. We study this failure mode with FLAT, a latent-conditioned reliability stress test for HFL backdoors. In FLAT, compromised clients still submit ordinary classifier updates to the server, while an attacker-side generator $G(x,t,z)$ separates target intent $t$ from trigger realization $z$. This separation shifts the audit question from whether one known trigger succeeds to how the hidden behavior varies across targets, latent samples, defenses, and post-stop rounds. On CIFAR-10, CIFAR-100, and Tiny-ImageNet, FLAT preserves clean utility while reaching 99.49%, 99.66%, and 94.10% single-target FedAvg ASR. The evaluation also reveals non-uniform defense responses, where a server rule can suppress one target mode while leaving another active. These observations motivate HFL backdoor audits that report target-wise ASR, worst-target ASR, target coverage, latent-sampled behavior, post-stop persistence, and defense response.

联邦学习后门攻击安全性隐变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。