arXiv:2606.31653cs.LGcs.AI2026-06

用教师模型的对抗信息提升认证鲁棒性,效果优于现有方法。

Improving Certified Robustness via Adversarial Distillation

  • 通过对抗蒸馏在logit空间传递教师模型的鲁棒特征,作为认证训练的下界替代
  • 在多个基准上实现当前最优的认证准确率,最高提升5.40个百分点
  • 适合关注模型认证鲁棒性且需兼顾标准准确率的研究者

认证训练旨在生成预测可形式化验证对抗扰动的模型,通常通过优化允许扰动集上最坏情况损失的上界来实现。基于紧松弛界的纯认证训练方法虽能保证可认证性,但会牺牲标准准确率;而对抗训练虽提升实证鲁棒性和标准准确率,却难以用神经网络验证器认证。近期研究表明,将对抗训练目标与基于区间边界传播(IBP)的松散上界结合,可在下界和上界之间实现有效插值。本文提出AD-CERT,一种融合对抗蒸馏与IBP上界的认证训练目标。结果表明,从一个具有实证鲁棒性的教师模型中,在logit空间蒸馏对抗信息,可作为认证训练的有效下界近似,使模型在多个鲁棒性基准上达到当前最优的认证性能。此外,在统一设置下,基于logit层面的对抗蒸馏相较鲁棒特征空间蒸馏,可提升最多5.40个百分点的认证准确率。

原文摘要 · Abstract (English)

Certified training aims to produce models whose predictions can be formally verified against adversarial perturbations, typically by optimising upper bounds on the worst-case loss over an allowed perturbation set. For neural networks, certified training methods based purely on tight relaxation bounds produce networks that are amenable to certification, but sacrifice standard accuracy. Conversely, adversarial training often yields stronger empirical robustness and standard accuracy, but the resulting models are generally difficult to certify with neural network verifiers. Recently, the literature has shown that better standard-certified accuracy trade-offs can be achieved by combining adversarial training objectives with loose over-approximations based on Interval Bound Propagation (IBP), effectively interpolating between lower and upper bounds of the worst-case loss. Building on this, we introduce AD-CERT, a certified training objective that combines adversarial distillation with an IBP upper bound. We show that distilling adversarial information over the logit space from an empirically robust teacher provides an effective lower bound surrogate for certified training, with AD-CERT achieving state-of-the-art certified performance on several robustness benchmarks. Furthermore, in a unified setup, distilling adversarial information at the logit-level is shown to improve certified accuracy over a robust feature-space distillation objective by up to 5.40 percentage points.

认证鲁棒性对抗蒸馏IBP模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。