arXiv:2603.29328cs.CRcs.AI2026-03中稿 · as a regular paper…被引 2

提出语义自然的后门攻击方法,让联邦学习更易受骗却不易被发现。

Beyond Corner Patches: Semantics-Aware Backdoor Attack in Federated Learning

  • 用真实语义内容(如戴眼镜)作触发器,避免突兀补丁。
  • 在多个数据集上实现高攻击成功率,且正常测试准确率不受影响。
  • 适合关注实际安全风险的研究者与系统设计者参考。

联邦学习中的后门攻击通常使用合成角落补丁或分布外模式评估,这些情况在现实中很少出现。本文在更贴近实际的设定下重新审视标准联邦学习(单个全局模型)的后门威胁:触发器需具备语义意义、在分布内且视觉合理。我们提出 SABLE,一种面向联邦学习的语义感知后门攻击方法,构建自然、内容一致的触发器(如改变佩戴太阳镜),并通过特征分离与参数正则化优化聚合感知的恶意目标,使攻击者更新接近良性更新。我们在 CelebA 发色分类和德国交通标志识别基准(GTSRB)上实现该方法,仅污染每个恶意客户端局部数据的一个小而可解释子集,其余遵循标准联邦学习协议。在异构客户端划分及多种聚合规则(FedAvg、Trimmed Mean、MultiKrum、FLAME)下,语义驱动的触发器均实现高目标攻击成功率,同时保持良性测试准确率。结果表明,语义对齐的后门仍是联邦学习中强大且现实的威胁,仅基于合成补丁触发器的鲁棒性声明可能过于乐观。

原文摘要 · Abstract (English)

Backdoor attacks on federated learning (FL) are most often evaluated with synthetic corner patches or out-of-distribution (OOD) patterns that are unlikely to arise in practice. In this paper, we revisit the backdoor threat to standard FL (a single global model) under a more realistic setting where triggers must be semantically meaningful, in-distribution, and visually plausible. We propose SABLE, a Semantics-Aware Backdoor for LEarning in federated settings, which constructs natural, content-consistent triggers (e.g., semantic attribute changes such as sunglasses) and optimizes an aggregation-aware malicious objective with feature separation and parameter regularization to keep attacker updates close to benign ones. We instantiate SABLE on CelebA hair-color classification and the German Traffic Sign Recognition Benchmark (GTSRB), poisoning only a small, interpretable subset of each malicious client's local data while otherwise following the standard FL protocol. Across heterogeneous client partitions and multiple aggregation rules (FedAvg, Trimmed Mean, MultiKrum, and FLAME), our semantics-driven triggers achieve high targeted attack success rates while preserving benign test accuracy. These results show that semantics-aligned backdoors remain a potent and practical threat in federated learning, and that robustness claims based solely on synthetic patch triggers can be overly optimistic.

联邦学习后门攻击语义触发安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。