对比扩散与非扩散防御模型,发现后者在迁移性和颜色泛化上更优。
Diffusion or Non-Diffusion Adversarial Defenses: Rethinking the Relation between Classifier and Adversarial Purifier
- 用分类器泛化损失分析扩散模型的防御机制
- 非扩散模型在CIFAR-10训练后直接用于ImageNet,性能超越专门训练的扩散模型
- 适合关注防御迁移性与少数据场景的研究者
对抗性防御研究仍面临先进攻击的挑战,而扩散模型正日益展现出防御潜力。不同于多数聚焦于测试时防御的先期研究,本文探讨了扩散模型引发的分类器泛化损失。通过对比基于扩散与非扩散的对抗净化器,发现非扩散模型在非自适应攻击的实际设定下同样表现优异。尽管非扩散模型已具良好鲁棒性,其在防御迁移性和颜色泛化方面尤为突出,且无需额外数据。值得注意的是,仅在CIFAR-10上训练的非扩散模型,在直接测试ImageNet时达到领先性能,超越在ImageNet上专门训练的现有扩散模型。
原文摘要 · Abstract (English)
Adversarial defense research continues to face challenges in combating against advanced adversarial attacks, yet with diffusion models increasingly favoring their defensive capabilities. Unlike most prior studies that focus on diffusion models for test-time defense, we explore the generalization loss in classifiers caused by diffusion models. We compare diffusion-based and non-diffusion-based adversarial purifiers, demonstrating that non-diffusion models can also achieve well performance under a practical setting of non-adaptive attack. While non-diffusion models show promising adversarial robustness, they particularly excel in defense transferability and color generalization without relying on additional data beyond the training set. Notably, a non-diffusion model trained on CIFAR-10 achieves state-of-the-art performance when tested directly on ImageNet, surpassing existing diffusion-based models trained specifically on ImageNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。