用强抗扰教师模型提升认证鲁棒模型性能。
Learning Better Certified Models from Empirically-Robust Teachers
- 用对抗训练的教师模型提取特征,指导学生模型学习。
- 在多个视觉基准上,认证鲁棒性与标准性能均优于当前最佳。
- 适合关注模型安全验证与实际性能平衡的研究者。
对抗训练通过在具体对抗扰动上训练,获得对特定攻击的强经验鲁棒性,但难以通过神经网络验证获得强鲁棒性证书。而早期的认证训练方法直接基于网络松弛的边界进行训练,虽能保证鲁棒性,但标准性能较差。近期研究通过结合对抗输出与神经网络边界的一族损失函数,实现了认证鲁棒性与标准性能之间的先进权衡。然而,与经验鲁棒性不同,可验证性仍显著损害标准性能。本文提出利用经验鲁棒的教师模型,通过知识蒸馏提升认证鲁棒模型的性能。采用通用的特征空间蒸馏目标,在一系列鲁棒计算机视觉基准上,针对ReLU网络,蒸馏结果一致超越现有认证训练方法的性能表现。
原文摘要 · Abstract (English)
Adversarial training attains strong empirical robustness to specific adversarial attacks by training on concrete adversarial perturbations, but it produces neural networks that are not amenable to strong robustness certificates through neural network verification. On the other hand, earlier certified training schemes directly train on bounds from network relaxations to obtain models that are certifiably robust, but display sub-par standard performance. Recent work has shown that state-of-the-art trade-offs between certified robustness and standard performance can be obtained through a family of losses combining adversarial outputs and neural network bounds. Nevertheless, differently from empirical robustness, verifiability still comes at a significant cost in standard performance. In this work, we propose to leverage empirically-robust teachers to improve the performance of certifiably-robust models through knowledge distillation. Using a versatile feature-space distillation objective, we show that distillation from adversarially-trained teachers consistently improves on the state-of-the-art in certified training for ReLU networks across a series of robust computer vision benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。