提出一种高效概率性鲁棒性认证方法,可适用于大型神经网络。
Probably Approximately Global Robustness Certification
- 通过采样ε-网并调用局部鲁棒性预言机,实现概率化全局鲁棒性验证。
- 样本量与输入维度、类别数和学习算法无关,突破传统方法瓶颈。
- 既比采样法更准确,又比形式验证法更可扩展,适合大规模模型。
我们提出并研究了分类算法的对抗鲁棒性的概率保证。传统形式验证方法难以计算,而基于采样的方法无法提供形式化保证;我们的方法能够高效地认证一种概率化的鲁棒性松弛。核心思想是采样一个ε-网,并在样本上调用局部鲁棒性预言机。值得注意的是,达到概率近似全局鲁棒性保证所需的样本量独立于输入维度、类别数量以及学习算法本身。因此,该方法可应用于超出传统形式验证范围的大规模神经网络。实验表明,其在刻画鲁棒性方面优于当前最先进的采样方法,且扩展性优于形式化方法。
原文摘要 · Abstract (English)
We propose and investigate probabilistic guarantees for the adversarial robustness of classification algorithms. While traditional formal verification approaches for robustness are intractable and sampling-based approaches do not provide formal guarantees, our approach is able to efficiently certify a probabilistic relaxation of robustness. The key idea is to sample an $ε$-net and invoke a local robustness oracle on the sample. Remarkably, the size of the sample needed to achieve probably approximately global robustness guarantees is independent of the input dimensionality, the number of classes, and the learning algorithm itself. Our approach can, therefore, be applied even to large neural networks that are beyond the scope of traditional formal verification. Experiments empirically confirm that it characterizes robustness better than state-of-the-art sampling-based approaches and scales better than formal methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。