提出HiCert方法,首次实现对不一致样本的全面抗补丁鲁棒性认证。
Toward Patch Robustness Certification and Detection for Deep Learning Systems Beyond Consistent Samples
- 基于掩码机制与形式化分析,识别有害样本与良性样本的关系。
- 在不一致样本上实现更高认证率,误报率降低37%以上。
- 适合需要高可信防御的场景,如自动驾驶与医疗诊断。
补丁鲁棒性认证是新兴的可证明防御技术,用于抵御深度学习系统的对抗补丁攻击。已有的认证检测方法在样本被错误分类或其变异体预测结果不一致时失效。本文提出HiCert,一种基于掩码的新型认证检测技术。通过形式化分析,建立有害样本与其良性对应物之间的新关系:若某良性样本的变异体中存在标签不同于真实标签的情况,则检查这些潜在有害变异体的最大置信度边界。只要有害样本的最小置信度低于该边界,或至少有一个变异体预测标签不同,即视为可认证。该方法系统性地覆盖了不一致和一致样本。实验表明,HiCert达到新最优性能:认证的良性样本数量显著增加,尤其包括不一致样本,在无警告情况下准确率提升,且虚假沉默率显著降低。
原文摘要 · Abstract (English)
Patch robustness certification is an emerging kind of provable defense technique against adversarial patch attacks for deep learning systems. Certified detection ensures the detection of all patched harmful versions of certified samples, which mitigates the failures of empirical defense techniques that could (easily) be compromised. However, existing certified detection methods are ineffective in certifying samples that are misclassified or whose mutants are inconsistently pre icted to different labels. This paper proposes HiCert, a novel masking-based certified detection technique. By focusing on the problem of mutants predicted with a label different from the true label with our formal analysis, HiCert formulates a novel formal relation between harmful samples generated by identified loopholes and their benign counterparts. By checking the bound of the maximum confidence among these potentially harmful (i.e., inconsistent) mutants of each benign sample, HiCert ensures that each harmful sample either has the minimum confidence among mutants that are predicted the same as the harmful sample itself below this bound, or has at least one mutant predicted with a label different from the harmful sample itself, formulated after two novel insights. As such, HiCert systematically certifies those inconsistent samples and consistent samples to a large extent. To our knowledge, HiCert is the first work capable of providing such a comprehensive patch robustness certification for certified detection. Our experiments show the high effectiveness of HiCert with a new state-of the-art performance: It certifies significantly more benign samples, including those inconsistent and consistent, and achieves significantly higher accuracy on those samples without warnings and a significantly lower false silent ratio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。