用扩散模型生成高逼真对抗补丁,让目标检测失效
AdvLogo: Adversarial Patch Attack against Object Detectors based on Diffusion Models
- 从语义空间出发,利用扩散模型在最后一层潜空间扰动生成对抗补丁
- 攻击成功率高且图像视觉质量好,未因分布偏移导致失真
- 适合研究对抗攻击、安全防御的人员,尤其关注视觉隐蔽性
随着深度学习快速发展,目标检测器表现优异,但在特定场景下仍存在漏洞。现有基于对抗补丁的研究常难以平衡攻击效果与图像质量。为此,我们提出一种从语义角度出发的新型补丁攻击框架——AdvLogo。基于假设:每个语义空间中均存在一个对抗子空间,使图像能诱导检测器误判,我们利用扩散模型去噪过程的语义理解能力,在最后一时刻对潜空间和无条件嵌入进行扰动,引导其进入对抗子区域。为缓解分布偏移对图像质量的负面影响,我们在频域中通过傅里叶变换对潜空间施加扰动。实验表明,AdvLogo在保持高视觉质量的同时实现了强大的攻击性能。
原文摘要 · Abstract (English)
With the rapid development of deep learning, object detectors have demonstrated impressive performance; however, vulnerabilities still exist in certain scenarios. Current research exploring the vulnerabilities using adversarial patches often struggles to balance the trade-off between attack effectiveness and visual quality. To address this problem, we propose a novel framework of patch attack from semantic perspective, which we refer to as AdvLogo. Based on the hypothesis that every semantic space contains an adversarial subspace where images can cause detectors to fail in recognizing objects, we leverage the semantic understanding of the diffusion denoising process and drive the process to adversarial subareas by perturbing the latent and unconditional embeddings at the last timestep. To mitigate the distribution shift that exposes a negative impact on image quality, we apply perturbation to the latent in frequency domain with the Fourier Transform. Experimental results demonstrate that AdvLogo achieves strong attack performance while maintaining high visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。