arXiv:2506.22506cs.CRcs.AI2025-06被引 2

提出防御联邦提示学习中的隐蔽后门攻击方法,提升模型安全性。

SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt Learning

  • 用嵌入空间异常检测器过滤恶意客户端的污染提示更新。
  • 在五个数据集上使后门攻击成功率降至0.8%,同时保持95%以上干净准确率。
  • 无需原始数据或标签,适用于多种场景,适合关注联邦学习安全的研究者。

联邦提示学习已成为一种高效且保护隐私的范式,用于在去中心化客户端上适配如CLIP的大规模视觉-语言模型。然而,该设置的安全性仍缺乏研究。本文首次研究了联邦提示学习中的后门攻击:当恶意客户端向输入图像注入视觉上难以察觉的可学习噪声触发器时,全局提示学习器会变得易受定向误分类影响,同时仍保持对干净输入的高准确率。针对此漏洞,我们提出了SABRE-FL,一种轻量、模块化的防御机制,通过离线训练的分布外数据嵌入空间异常检测器过滤被污染的提示更新。SABRE-FL无需访问原始客户端数据或标签,并在多种数据集上具有泛化能力。理论与实证均表明,恶意客户端可被可靠识别并过滤。在五个多样化数据集和四种基线防御下,SABRE-FL显著降低后门准确率,同时保持干净准确率,展现出强大性能,凸显未来联邦系统中鲁棒提示学习的重要性。

原文摘要 · Abstract (English)

Federated Prompt Learning has emerged as a communication-efficient and privacy-preserving paradigm for adapting large vision-language models like CLIP across decentralized clients. However, the security implications of this setup remain underexplored. In this work, we present the first study of backdoor attacks in Federated Prompt Learning. We show that when malicious clients inject visually imperceptible, learnable noise triggers into input images, the global prompt learner becomes vulnerable to targeted misclassification while still maintaining high accuracy on clean inputs. Motivated by this vulnerability, we propose SABRE-FL, a lightweight, modular defense that filters poisoned prompt updates using an embedding-space anomaly detector trained offline on out-of-distribution data. SABRE-FL requires no access to raw client data or labels and generalizes across diverse datasets. We show, both theoretically and empirically, that malicious clients can be reliably identified and filtered using an embedding-based detector. Across five diverse datasets and four baseline defenses, SABRE-FL outperforms all baselines by significantly reducing backdoor accuracy while preserving clean accuracy, demonstrating strong empirical performance and underscoring the need for robust prompt learning in future federated systems.

联邦学习后门攻击提示学习安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。