用低秩迭代扩散法,高效清除对抗样本中的恶意扰动。
LoRID: Low-Rank Iterative Diffusion for Adversarial Purification
- 多轮早期扩散去噪结合泰克分解,逐次净化噪声。
- 在CIFAR-10/100、CelebA-HQ、ImageNet上实现更强鲁棒性。
- 适用于白盒与灰盒攻击场景,净化误差更低。
本文从信息论角度分析基于扩散模型的净化方法,这类方法是当前最先进的对抗防御技术,能通过扩散模型去除对抗样本中的恶意扰动。通过理论刻画基于马尔可夫过程的扩散净化固有误差,我们提出一种新型低秩迭代扩散净化方法——LoRID,旨在以低内在净化误差去除对抗扰动。LoRID采用多阶段净化流程,在扩散模型的早期时间步进行多轮扩散-去噪循环,并引入张量分解(Tucker decomposition)这一矩阵分解的扩展形式,以在高噪声环境下有效去除对抗噪声。该方法显著增加有效扩散时间步数,可抵御强对抗攻击,在CIFAR-10/100、CelebA-HQ和ImageNet数据集上均展现出优越的鲁棒性表现,适用于白盒与灰盒攻击场景。
原文摘要 · Abstract (English)
This work presents an information-theoretic examination of diffusion-based purification methods, the state-of-the-art adversarial defenses that utilize diffusion models to remove malicious perturbations in adversarial examples. By theoretically characterizing the inherent purification errors associated with the Markov-based diffusion purifications, we introduce LoRID, a novel Low-Rank Iterative Diffusion purification method designed to remove adversarial perturbation with low intrinsic purification errors. LoRID centers around a multi-stage purification process that leverages multiple rounds of diffusion-denoising loops at the early time-steps of the diffusion models, and the integration of Tucker decomposition, an extension of matrix factorization, to remove adversarial noise at high-noise regimes. Consequently, LoRID increases the effective diffusion time-steps and overcomes strong adversarial attacks, achieving superior robustness performance in CIFAR-10/100, CelebA-HQ, and ImageNet datasets under both white-box and black-box settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。