用动量优化提升扩散模型生成的对抗样本隐蔽性与攻击效果
Boosting Imperceptibility of Stable Diffusion-based Adversarial Examples Generation with Momentum
- 通过调整文本嵌入引导扩散模型生成对抗图,保持原图语义
- 生成对抗样本误分类率79%,较现有方法提升35%
- 适合用于评估图像分类器的鲁棒性,兼顾隐蔽性与攻击力
我们提出一种新框架SD-MIAE,基于Stable Diffusion生成能有效误导神经网络分类器的对抗样本,同时保持视觉不可察觉性和与原始类别标签的语义相似性。该方法通过操纵Stable Diffusion模型中指定类别的文本嵌入,在其潜在空间中引导对抗图像生成,确保图像具有高视觉保真度。框架包含两个阶段:(1) 初始对抗优化阶段,修改文本嵌入以生成被错误分类但外观自然的图像;(2) 动量优化阶段,进一步精炼对抗扰动。引入动量机制使扰动在迭代中更稳定,提升了误分类率和图像保真度。实验表明,SD-MIAE实现79%的高误分类率,相比当前最优方法提升35%,同时保持对抗扰动的不可察觉性及原始类别的语义一致性,是一种实用的鲁棒性评估方法。
原文摘要 · Abstract (English)
We propose a novel framework, Stable Diffusion-based Momentum Integrated Adversarial Examples (SD-MIAE), for generating adversarial examples that can effectively mislead neural network classifiers while maintaining visual imperceptibility and preserving the semantic similarity to the original class label. Our method leverages the text-to-image generation capabilities of the Stable Diffusion model by manipulating token embeddings corresponding to the specified class in its latent space. These token embeddings guide the generation of adversarial images that maintain high visual fidelity. The SD-MIAE framework consists of two phases: (1) an initial adversarial optimization phase that modifies token embeddings to produce misclassified yet natural-looking images and (2) a momentum-based optimization phase that refines the adversarial perturbations. By introducing momentum, our approach stabilizes the optimization of perturbations across iterations, enhancing both the misclassification rate and visual fidelity of the generated adversarial examples. Experimental results demonstrate that SD-MIAE achieves a high misclassification rate of 79%, improving by 35% over the state-of-the-art method while preserving the imperceptibility of adversarial perturbations and the semantic similarity to the original class label, making it a practical method for robust adversarial evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。