针对生成数据增强的隐形后门攻击,提升成功率46.43%。
Invisible Clean-Label Backdoor Attacks for Generative Data Augmentation
- 在隐空间引入扰动实现隐形后门,突破像素级攻击局限。
- 平均攻击成功率提升46.43%,清洁准确率几乎不变。
- 对现有防御方法高度鲁棒,适合研究生成模型安全者。
随着图像生成模型的快速发展,生成式数据增强已成为在小规模数据集下丰富训练样本的有效手段。然而,在实际应用中,该方法可能面临干净标签后门攻击,此类攻击旨在绕过人工检测。基于理论分析与初步实验,我们发现直接将现有像素级干净标签后门攻击方法(如COMBAT)应用于生成图像时,攻击成功率较低。这促使我们超越像素级触发机制,转而关注隐空间特征层面。为此,我们提出InvLBA——一种基于隐空间扰动的生成数据增强隐形干净标签后门攻击方法。理论上证明了其清洁准确率与攻击成功率的泛化能力可被保证。在多个数据集上的实验表明,本方法平均攻击成功率提升46.43%,同时清洁准确率几乎无损,并对当前最先进的防御方法表现出强鲁棒性。
原文摘要 · Abstract (English)
With the rapid advancement of image generative models, generative data augmentation has become an effective way to enrich training images, especially when only small-scale datasets are available. At the same time, in practical applications, generative data augmentation can be vulnerable to clean-label backdoor attacks, which aim to bypass human inspection. However, based on theoretical analysis and preliminary experiments, we observe that directly applying existing pixel-level clean-label backdoor attack methods (e.g., COMBAT) to generated images results in low attack success rates. This motivates us to move beyond pixel-level triggers and focus instead on the latent feature level. To this end, we propose InvLBA, an invisible clean-label backdoor attack method for generative data augmentation by latent perturbation. We theoretically prove that the generalization of the clean accuracy and attack success rates of InvLBA can be guaranteed. Experiments on multiple datasets show that our method improves the attack success rate by 46.43% on average, with almost no reduction in clean accuracy and high robustness against SOTA defense methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。