arXiv:2411.14834cs.LGcs.CR2024-11被引 4

该防御看似鲁棒,实则可被自适应攻击大幅突破。

Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense

  • 通过多分辨率噪声融合中间表征实现防御
  • 在CIFAR-10上鲁棒准确率从62%降至11%
  • 适合研究对抗样本防御漏洞的学者

最近提出的'全局集成'防御方法通过在多个带噪图像分辨率下集成模型中间表示,声称能提升图像分类器对对抗样本的鲁棒性。该方法曾显示对多种先进攻击有效,并且模型梯度具有感知一致性:攻击产生的噪声与目标类别视觉相似。本文指出该防御并不真正鲁棒。我们首先揭示其随机性和集成机制导致严重梯度掩蔽。随后采用标准自适应攻击,在ℓ∞范数威胁模型下(ε=8/255),使CIFAR-100上的鲁棒准确率从48%降至14%,CIFAR-10上的从62%降至11%。

原文摘要 · Abstract (English)

Ensemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations at multiple noisy image resolutions, producing a single robust classification. This defense was shown to be effective against multiple state-of-the-art attacks. Perhaps even more convincingly, it was shown that the model's gradients are perceptually aligned: attacks against the model produce noise that perceptually resembles the targeted class. In this short note, we show that this defense is not robust to adversarial attack. We first show that the defense's randomness and ensembling method cause severe gradient masking. We then use standard adaptive attack techniques to reduce the defense's robust accuracy from 48% to 14% on CIFAR-100 and from 62% to 11% on CIFAR-10, under the $\ell_\infty$-norm threat model with $\varepsilon=8/255$.

对抗防御梯度掩蔽鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。