通过优化重建误差的平坦区域,提升对抗样本净化后的鲁棒性。
PurSAMERE: Reliable Adversarial Purification via Sharpness-Aware Minimization of Expected Reconstruction Error
- 基于期望重建误差最小化,在局部搜索净化样本。
- 在强白盒攻击下,相比现有方法提升显著鲁棒性。
- 适合追求可靠且稳定对抗防御的应用场景。
我们提出一种新的确定性净化方法,通过将潜在对抗样本映射到靠近数据分布模式的邻近点,从而提升分类器的可靠性。该方法利用噪声损坏数据的期望重建误差进行训练,学习输入数据分布的结构特征。给定一个潜在对抗输入,方法在其局部邻域内搜索在噪声扰动下期望重建误差最小的净化样本,并将其输入分类器。净化过程中采用尖锐度感知最小化,引导净化样本趋向期望重建误差曲面的平坦区域,从而增强鲁棒性。我们进一步证明,当噪声水平降低时,最小化期望重建误差会使净化样本偏向高斯平滑密度的局部极大值;在评分模型的局部假设下,可证明在小噪声极限下恢复局部极大值。实验表明,在强确定性白盒攻击下,该方法显著优于现有最先进方法。
原文摘要 · Abstract (English)
We propose a novel deterministic purification method to improve adversarial robustness by mapping a potentially adversarial sample toward a nearby sample that lies close to a mode of the data distribution, where classifiers are more reliable. We design the method to be deterministic to ensure reliable test accuracy and to prevent the degradation of effective robustness observed in stochastic purification approaches when the adversary has full knowledge of the system and its randomness. We employ a score model trained by minimizing the expected reconstruction error of noise-corrupted data, thereby learning the structural characteristics of the input data distribution. Given a potentially adversarial input, the method searches within its local neighborhood for a purified sample that minimizes the expected reconstruction error under noise corruption and then feeds this purified sample to the classifier. During purification, sharpness-aware minimization is used to guide the purified samples toward flat regions of the expected reconstruction error landscape, thereby enhancing robustness. We further show that, as the noise level decreases, minimizing the expected reconstruction error biases the purified sample toward local maximizers of the Gaussian-smoothed density; under additional local assumptions on the score model, we prove recovery of a local maximizer in the small-noise limit. Experimental results demonstrate significant gains in adversarial robustness over state-of-the-art methods under strong deterministic white-box attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。