基于注意力机制生成最优触发器,提升后门攻击效果与隐蔽性。
An Effective and Resilient Backdoor Attack Framework against Deep Neural Networks and Vision Transformers
- 用注意力机制搜索最佳触发器形状与位置
- 低污染比例下攻击成功率提升82%,样本自然度高
- 适用于DNN与视觉变换模型,可绕过主流防御
近期研究揭示了深度神经网络(DNN)易受后门攻击。然而,现有方法随意设定触发器掩码或随机选取触发器,限制了触发器的有效性与鲁棒性。本文提出一种基于注意力的掩码生成方法,自动搜索最优触发器形状与位置,并在损失函数中引入用户体验质量(QoE)项,精细调节触发器透明度,使被污染样本更自然。为提升目标模型预测准确率,提出交替重训练算法:偶数轮使用混合中毒数据集,奇数轮仅用正常样本。此外,在联合优化框架下交替优化触发器与中毒模型,进一步提升攻击性能。该方法扩展至视觉变换模型。在VGG-Flower、CIFAR-10、GTSRB、CIFAR-100和ImageNette数据集上的实验表明,当毒化比例低时,攻击成功率较基线最高提升82%,且被污染样本保持高QoE。所提框架对当前主流后门防御也表现出强鲁棒性。
原文摘要 · Abstract (English)
Recent studies have revealed the vulnerability of Deep Neural Network (DNN) models to backdoor attacks. However, existing backdoor attacks arbitrarily set the trigger mask or use a randomly selected trigger, which restricts the effectiveness and robustness of the generated backdoor triggers. In this paper, we propose a novel attention-based mask generation methodology that searches for the optimal trigger shape and location. We also introduce a Quality-of-Experience (QoE) term into the loss function and carefully adjust the transparency value of the trigger in order to make the backdoored samples to be more natural. To further improve the prediction accuracy of the victim model, we propose an alternating retraining algorithm in the backdoor injection process. The victim model is retrained with mixed poisoned datasets in even iterations and with only benign samples in odd iterations. Besides, we launch the backdoor attack under a co-optimized attack framework that alternately optimizes the backdoor trigger and backdoored model to further improve the attack performance. Apart from DNN models, we also extend our proposed attack method against vision transformers. We evaluate our proposed method with extensive experiments on VGG-Flower, CIFAR-10, GTSRB, CIFAR-100, and ImageNette datasets. It is shown that we can increase the attack success rate by as much as 82\% over baselines when the poison ratio is low and achieve a high QoE of the backdoored samples. Our proposed backdoor attack framework also showcases robustness against state-of-the-art backdoor defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。