通过操控特征嵌入实现低毒率、高隐蔽的后门攻击
Poison in the Well: Feature Embedding Disruption in Backdoor Attacks
- 针对神经网络特征空间,设计聚类优化策略提升攻击稳定性
- 毒化率低至0.01%时仍达100%攻击成功率,正常准确率衰减≤1%
- 适用于清洁标签与脏标签场景,适合研究防御机制的学者参考
后门攻击通过在训练数据中嵌入恶意触发器,使神经网络在推理阶段被操控,同时保持对良性输入的高准确率。然而,现有方法存在对训练数据依赖过高、隐蔽性差、不稳定等问题,限制了其在真实场景的应用。本文提出ShadowPrint,一种针对神经网络特征嵌入的通用后门攻击方法,显著降低对训练数据的依赖,在极低毒化率(低至0.01%)下仍有效。该方法采用基于聚类的优化策略对齐特征嵌入,确保在多种场景下的鲁棒性。大量实验表明,ShadowPrint在干净标签和脏标签设置下均实现高达100%的攻击成功率(ASR),正常准确率(CA)衰减不超过1%,毒化检测率(DDR)平均低于5%,毒化率范围为0.01%至0.05%,树立了后门攻击的新标准,凸显针对特征空间操纵的高级防御策略的必要性。
原文摘要 · Abstract (English)
Backdoor attacks embed malicious triggers into training data, enabling attackers to manipulate neural network behavior during inference while maintaining high accuracy on benign inputs. However, existing backdoor attacks face limitations manifesting in excessive reliance on training data, poor stealth, and instability, which hinder their effectiveness in real-world applications. Therefore, this paper introduces ShadowPrint, a versatile backdoor attack that targets feature embeddings within neural networks to achieve high ASRs and stealthiness. Unlike traditional approaches, ShadowPrint reduces reliance on training data access and operates effectively with exceedingly low poison rates (as low as 0.01%). It leverages a clustering-based optimization strategy to align feature embeddings, ensuring robust performance across diverse scenarios while maintaining stability and stealth. Extensive evaluations demonstrate that ShadowPrint achieves superior ASR (up to 100%), steady CA (with decay no more than 1% in most cases), and low DDR (averaging below 5%) across both clean-label and dirty-label settings, and with poison rates ranging from as low as 0.01% to 0.05%, setting a new standard for backdoor attack capabilities and emphasizing the need for advanced defense strategies focused on feature space manipulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。