发现去噪模型易受跨模型攻击,提出基于典型集采样的新防御方法。
Adversarial Transferability in Deep Denoising Models: Theoretical Insights and Robustness Enhancement via Out-of-Distribution Typical Set Sampling
- 用典型集理论解释噪声与对抗样本的分布差异
- 新方法在多种去噪模型上显著提升抗攻击能力
- 适合关注模型鲁棒性的图像处理研究者
基于深度学习的图像去噪模型表现优异,但其鲁棒性分析不足。一个关键问题是,这些模型容易受到对抗攻击——微小且精心设计的扰动即可导致模型失效。令人惊讶的是,针对某一模型设计的扰动可轻易转移至其他多种模型(包括CNN、Transformer、展开模型及即插即用模型),引发广泛失败,而分类模型中并未观察到此现象。本文通过一系列假设与实验,分析高对抗迁移性的潜在原因。利用典型集与渐近等分性质,证明对抗样本仅轻微偏离原始输入分布的典型集,从而导致模型失效。基于此,提出一种新型防御策略:离群典型集采样训练(TS)。该方法不仅显著增强模型鲁棒性,相比原模型还小幅提升去噪性能。
原文摘要 · Abstract (English)
Deep learning-based image denoising models demonstrate remarkable performance, but their lack of robustness analysis remains a significant concern. A major issue is that these models are susceptible to adversarial attacks, where small, carefully crafted perturbations to input data can cause them to fail. Surprisingly, perturbations specifically crafted for one model can easily transfer across various models, including CNNs, Transformers, unfolding models, and plug-and-play models, leading to failures in those models as well. Such high adversarial transferability is not observed in classification models. We analyze the possible underlying reasons behind the high adversarial transferability through a series of hypotheses and validation experiments. By characterizing the manifolds of Gaussian noise and adversarial perturbations using the concept of typical set and the asymptotic equipartition property, we prove that adversarial samples deviate slightly from the typical set of the original input distribution, causing the models to fail. Based on these insights, we propose a novel adversarial defense method: the Out-of-Distribution Typical Set Sampling Training strategy (TS). TS not only significantly enhances the model's robustness but also marginally improves denoising performance compared to the original model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。