用低秩方法防御扩散模型的对抗攻击,提升安全性。
A Low-Rank Defense Method for Adversarial Attack on Diffusion Models
- 引入低秩适配模块与平衡参数检测对抗样本。
- 在人脸和风景图像上防御效果优于基线方法。
- 适合需要安全生成高质量图像的应用场景。
近期,针对扩散模型及其微调过程的对抗攻击迅速发展。为防止这些攻击算法影响扩散模型的实际应用,亟需相应的防御策略。本文提出一种高效的防御方法——低秩防御(LoRD),用于防御潜在扩散模型(LDMs)的对抗攻击。LoRD结合低秩适配(LoRA)模块、融合思想与平衡参数,实现对抗样本的检测与防御。基于此,我们构建了防御流程,将学习到的LoRD模块应用于扩散模型以抵御攻击算法。实验表明,经洛德方法微调后的LDM在同时包含对抗样本和干净样本的数据集上,仍可生成高质量图像。我们在人脸与风景图像上进行了广泛实验,结果表明该方法在防御性能上显著优于基线方法。
原文摘要 · Abstract (English)
Recently, adversarial attacks for diffusion models as well as their fine-tuning process have been developed rapidly. To prevent the abuse of these attack algorithms from affecting the practical application of diffusion models, it is critical to develop corresponding defensive strategies. In this work, we propose an efficient defensive strategy, named Low-Rank Defense (LoRD), to defend the adversarial attack on Latent Diffusion Models (LDMs). LoRD introduces the merging idea and a balance parameter, combined with the low-rank adaptation (LoRA) modules, to detect and defend the adversarial samples. Based on LoRD, we build up a defense pipeline that applies the learned LoRD modules to help diffusion models defend against attack algorithms. Our method ensures that the LDM fine-tuned on both adversarial and clean samples can still generate high-quality images. To demonstrate the effectiveness of our approach, we conduct extensive experiments on facial and landscape images, and our method shows significantly better defense performance compared to the baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。