触发器颜色影响联邦后门攻击成功率,白触发动态更易攻陷金发类别。
Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks

- 用不同颜色的视觉触发器(如口罩、墨镜)实施语义驱动后门攻击。
- 白色触发器对金发类别攻击成功率更高,黑色对黑发类别更有效。
- 适用于研究联邦学习安全性的研究人员与防御设计者。
联邦学习在恶意客户端注入中毒更新时存在后门攻击风险,同时保持良性任务性能。本文研究一种语义驱动的后门机制,攻击者使用自然视觉配件作为触发器,仅改变触发器颜色而固定攻击流程。实验采用面具、太阳镜等语义触发物,分别以黑白变体实例化,在四分类的CelebA发色任务中评估其效果。恶意客户端通过在源类别图像上添加触发器并重标为目标类别构建中毒样本,良性客户端仅训练干净数据。攻击在标准中毒目标和更强的SABLE目标下测试,后者结合了分类损失、触发目标损失、倒数第二层特征分离损失及正则项,以减少更新漂移。实验表明,即使触发语义、位置和中毒预算不变,触发器颜色显著影响攻击成功率:白色触发器更利于攻击金发类别,黑色触发器对黑发类别更有效。该趋势在鲁棒聚合下依然成立,证明触发器颜色是联邦学习中语义后门机制运作与评估的关键因素。
原文摘要 · Abstract (English)
Federated learning is vulnerable to backdoor attacks in which malicious clients inject poisoned updates while preserving benign-task performance. In this paper, we study a semantics-driven backdoor mechanism in which attackers use natural visual accessories as triggers and manipulate only the trigger color while keeping the attack pipeline fixed. Our framework considers semantic trigger objects such as masks and sunglasses, instantiated in black and white variants, and evaluates their effect in a controlled federated learning setting. Malicious clients construct poisoned samples by applying a trigger to source-class images and relabeling them to an attacker-chosen target class, while benign clients train only on clean data. We analyze this mechanism under both a standard poisoning objective and a stronger SABLE-based objective that combines clean classification loss, triggered target loss, feature-separation loss in the penultimate representation space, and regularization to keep malicious updates close to the global model. This design enables the attack to remain effective while reducing excessive update drift. Experiments on a four-class CelebA hair-color task show that trigger color significantly changes attack success rate even when trigger semantics, placement, and poisoning budget are unchanged. White triggers are more effective for attacks targeting the blond class, whereas black triggers perform better for attacks targeting the black class. The same trend persists under robust aggregation, showing that trigger color is a meaningful factor in the operation, persistence, and evaluation of semantic backdoor mechanisms in federated learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。