用语义优化生成自然且隐蔽的对抗样本,提升攻击成功率。
SemDiff: Generating Natural Unrestricted Adversarial Examples via Semantic Attributes Optimization in Diffusion Models
- 在扩散模型语义空间中优化多属性,实现精准控制
- 在三大高分辨率数据集上攻击成功率超现有方法
- 生成样本自然且能绕过多种防御机制,适合安全评估
无限制对抗样本(UAEs)允许攻击者在无需原始样本的情况下生成不受约束的对抗样本,对深度学习模型安全构成严重威胁。现有方法利用扩散模型生成UAEs,但通常因仅在中间潜在噪声上优化,导致生成样本缺乏自然性和隐蔽性。为此,我们提出SemDiff,一种新型无限制对抗攻击方法,通过探索扩散模型的语义潜在空间,设计多属性优化策略,在确保攻击成功的同时保持生成样本的自然性和隐蔽性。我们在三个高分辨率数据集(CelebA-HQ、AFHQ、ImageNet)上的四个任务上进行了广泛实验。结果表明,SemDiff在攻击成功率和隐蔽性方面均优于现有最优方法。生成的UAEs具有自然外观且呈现语义上有意义的变化,与属性权重一致。此外,SemDiff可有效规避多种防御机制,进一步验证其有效性与威胁性。
原文摘要 · Abstract (English)
Unrestricted adversarial examples (UAEs), allow the attacker to create non-constrained adversarial examples without given clean samples, posing a severe threat to the safety of deep learning models. Recent works utilize diffusion models to generate UAEs. However, these UAEs often lack naturalness and imperceptibility due to simply optimizing in intermediate latent noises. In light of this, we propose SemDiff, a novel unrestricted adversarial attack that explores the semantic latent space of diffusion models for meaningful attributes, and devises a multi-attributes optimization approach to ensure attack success while maintaining the naturalness and imperceptibility of generated UAEs. We perform extensive experiments on four tasks on three high-resolution datasets, including CelebA-HQ, AFHQ and ImageNet. The results demonstrate that SemDiff outperforms state-of-the-art methods in terms of attack success rate and imperceptibility. The generated UAEs are natural and exhibit semantically meaningful changes, in accord with the attributes' weights. In addition, SemDiff is found capable of evading different defenses, which further validates its effectiveness and threatening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。