用认证训练提升单步攻击下的鲁棒性,缓解灾难性过拟合。
On Using Certified Training towards Empirical Robustness
- 结合对抗攻击与网络超逼近,改进认证训练以抵抗单步攻击。
- 在合适设置下,性能接近多步攻击基线,缩小与认证防御的差距。
- 提出简单正则化方法,显著降低训练时间,适合实际部署。
对抗训练是提升模型对特定对抗样本鲁棒性的主流方法。基于多步攻击的变体计算开销大,而单步变体易受灾难性过拟合影响,限制了其在大扰动下的实用性。另一条研究路径——认证训练——旨在为任何可能攻击提供形式化鲁棒性保证,但当前最优的实证防御与认证防御之间存在巨大性能差距,严重限制了后者的应用。受近期认证训练中结合对抗攻击与网络超逼近的启发,并注意到局部线性与灾难性过拟合之间的关联,本文实验验证了使用认证训练提升实证鲁棒性的潜力与局限。结果表明,经过适当调优,近期的认证训练算法可有效防止单步攻击下的灾难性过拟合;在合适的实验设置下,其性能可逼近多步攻击基线。最后,我们提出一种概念简单的网络超逼近正则化方法,实现类似效果的同时显著降低运行时间。
原文摘要 · Abstract (English)
Adversarial training is arguably the most popular way to provide empirical robustness against specific adversarial examples. While variants based on multi-step attacks incur significant computational overhead, single-step variants are vulnerable to a failure mode known as catastrophic overfitting, which hinders their practical utility for large perturbations. A parallel line of work, certified training, has focused on producing networks amenable to formal guarantees of robustness against any possible attack. However, the wide gap between the best-performing empirical and certified defenses has severely limited the applicability of the latter. Inspired by recent developments in certified training, which rely on a combination of adversarial attacks with network over-approximations, and by the connections between local linearity and catastrophic overfitting, we present experimental evidence on the practical utility and limitations of using certified training towards empirical robustness. We show that, when tuned for the purpose, a recent certified training algorithm can prevent catastrophic overfitting on single-step attacks, and that it can bridge the gap to multi-step baselines under appropriate experimental settings. Finally, we present a conceptually simple regularizer for network over-approximations that can achieve similar effects while markedly reducing runtime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。