用扩散模型的特征表示提升对抗训练鲁棒性
Expanding the Role of Diffusion Models for Robust Classifier Training
- 将扩散模型的内部特征作为辅助信号加入对抗训练
- 在多个数据集上显著提升分类器的对抗鲁棒性
- 适合关注模型安全性与表征学习的研究者
将扩散生成的合成数据融入对抗训练(AT)已被证明可显著提升鲁棒图像分类器的训练效果。本文拓展了扩散模型的作用,不仅用于生成数据,还探索其内部表征是否能带来额外增益。系统实验表明,扩散模型的表征具有多样性和部分鲁棒性,将其作为辅助学习信号引入AT能持续提升各类设置下的鲁棒性。进一步分析显示,扩散模型的引入促使特征更解耦;同时,扩散表征与生成数据在塑造表征方面发挥互补作用。在CIFAR-10、CIFAR-100和ImageNet上的实验验证了该方法的有效性,表明联合利用扩散表征与合成数据在对抗训练中具有显著优势。
原文摘要 · Abstract (English)
Incorporating diffusion-generated synthetic data into adversarial training (AT) has been shown to substantially improve the training of robust image classifiers. In this work, we extend the role of diffusion models beyond merely generating synthetic data, examining whether their internal representations, which encode meaningful features of the data, can provide additional benefits for robust classifier training. Through systematic experiments, we show that diffusion models offer representations that are both diverse and partially robust, and that explicitly incorporating diffusion representations as an auxiliary learning signal during AT consistently improves robustness across settings. Furthermore, our representation analysis indicates that incorporating diffusion models into AT encourages more disentangled features, while diffusion representations and diffusion-generated synthetic data play complementary roles in shaping representations. Experiments on CIFAR-10, CIFAR-100, and ImageNet validate these findings, demonstrating the effectiveness of jointly leveraging diffusion representations and synthetic data within AT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。