通过神经元重要性排序提升模型抗攻击能力,不需训练且效率极高。
Interpretability-Guided Test-Time Adversarial Defense

- 基于可解释性识别关键神经元,动态调整输出以增强鲁棒性。
- 在CIFAR10/100和ImageNet上分别提升2.6%/4.9%/2.8%准确率。
- 对黑盒、白盒及自适应攻击均有效,比现有方法快4倍。
我们提出一种新颖且低成本的测试时对抗防御方法,通过设计基于可解释性的神经元重要性排序机制,识别对输出类别重要的神经元。该方法为无需训练的防御策略,在显著提升鲁棒性-准确率权衡的同时,计算开销极小。作为最高效的测试时防御之一(速度提升4倍),该方法对多种黑盒、白盒及自适应攻击均具备强鲁棒性,能突破此前测试时防御的局限。我们在标准RobustBench基准上验证了其在CIFAR10、CIFAR100和ImageNet-1k上的有效性,平均准确率提升分别为2.6%、4.9%和2.8%。即使在强自适应攻击下,仍比当前最优测试时防御平均提升1.5%。
原文摘要 · Abstract (English)
We propose a novel and low-cost test-time adversarial defense by devising interpretability-guided neuron importance ranking methods to identify neurons important to the output classes. Our method is a training-free approach that can significantly improve the robustness-accuracy tradeoff while incurring minimal computational overhead. While being among the most efficient test-time defenses (4x faster), our method is also robust to a wide range of black-box, white-box, and adaptive attacks that break previous test-time defenses. We demonstrate the efficacy of our method for CIFAR10, CIFAR100, and ImageNet-1k on the standard RobustBench benchmark (with average gains of 2.6%, 4.9%, and 2.8% respectively). We also show improvements (average 1.5%) over the state-of-the-art test-time defenses even under strong adaptive attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。