将不确定性量化融入随机平滑,提升模型抗攻击鲁棒性
Integrating uncertainty quantification into randomized smoothing based robustness guarantees
- 用高不确定性时拒识的分类器融合随机平滑框架
- 在CIFAR10上鲁棒半径比传统模型大20.93%
- 适合关注安全性的模型开发者与不确定性研究者
深度神经网络虽强大,却易受对抗攻击影响,危及关键应用。随机平滑提供概率性保证:输入周围ℓ₂球内预测不变。而基于不确定性的拒识策略常用于防御攻击。本文将两者结合,构建可拒识高不确定性输入的分类器,并推导出两类新鲁棒性保障:(i)预测标签不变且不确定性低的ℓ₂球半径;(ii)预测不改变或进入不确定状态的ℓ₂球半径。前者防范引发不确定性的攻击,后者揭示导致错误预测所需扰动量。在CIFAR10上,该框架相比无拒识机制模型,鲁棒半径提升达20.93%。实验表明,该框架可系统评估不同网络结构与不确定性度量,识别理想不确定性量化特性,同时增强分布外检测能力。
原文摘要 · Abstract (English)
Deep neural networks have proven to be extremely powerful, however, they are also vulnerable to adversarial attacks which can cause hazardous incorrect predictions in safety-critical applications. Certified robustness via randomized smoothing gives a probabilistic guarantee that the smoothed classifier's predictions will not change within an $\ell_2$-ball around a given input. On the other hand (uncertainty) score-based rejection is a technique often applied in practice to defend models against adversarial attacks. In this work, we fuse these two approaches by integrating a classifier that abstains from predicting when uncertainty is high into the certified robustness framework. This allows us to derive two novel robustness guarantees for uncertainty aware classifiers, namely (i) the radius of an $\ell_2$-ball around the input in which the same label is predicted and uncertainty remains low and (ii) the $\ell_2$-radius of a ball in which the predictions will either not change or be uncertain. While the former provides robustness guarantees with respect to attacks aiming at increased uncertainty, the latter informs about the amount of input perturbation necessary to lead the uncertainty aware model into a wrong prediction. Notably, this is on CIFAR10 up to 20.93% larger than for models not allowing for uncertainty based rejection. We demonstrate, that the novel framework allows for a systematic robustness evaluation of different network architectures and uncertainty measures and to identify desired properties of uncertainty quantification techniques. Moreover, we show that leveraging uncertainty in a smoothed classifier helps out-of-distribution detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。