为训练和测试阶段的攻击提供可证明的鲁棒性保障
Robustness Certificates for Neural Networks Against Data Poisoning and Evasion Attacks
- 将梯度训练建模为离散动力系统,用屏障函数验证鲁棒性
- 在多个数据集上认证出非平凡的扰动容忍范围
- 首个统一框架,同时覆盖训练与测试时攻击防御
机器学习在安全关键领域的广泛应用加剧了对抗性威胁的风险,尤其是通过污染训练数据来降低性能或引发不安全行为的数据投毒攻击。现有防御方法大多缺乏形式化保证,或依赖于对模型类别、攻击类型、污染程度的限制性假设,且多基于逐点认证,实用性受限。本文提出一种基于形式化鲁棒性认证的原理性框架,将梯度训练建模为离散时间动力系统(dt-DS),并将投毒鲁棒性问题转化为形式化安全验证问题。通过借鉴控制理论中的屏障证书(BCs)概念,我们引入充分条件,以确保在最坏情况下的ℓ_p范数投毒下,最终模型仍保持安全。为实现实际应用,我们将屏障证书参数化为神经网络,并在有限的污染轨迹集合上进行训练。进一步通过求解场景凸规划(SCP)推导出大概率正确(PAC)边界,从而获得超越训练集的认证鲁棒半径置信下界。重要的是,该框架还扩展至测试时攻击的认证,成为首个在训练与测试时攻击场景中均提供形式化保证的统一框架。在MNIST、SVHN、CIFAR-10和CIFAR-100上的实验表明,该方法能在无需攻击先验知识或污染水平信息的情况下,实现模型无关的非平凡扰动预算认证。
原文摘要 · Abstract (English)
The increasing use of machine learning in safety-critical domains amplifies the risk of adversarial threats, especially data poisoning attacks that corrupt training data to degrade performance or induce unsafe behavior. Most existing defenses lack formal guarantees or rely on restrictive assumptions about the model class, attack type, extent of poisoning, or point-wise certification, limiting their practical reliability. This paper introduces a principled formal robustness certification framework that models gradient-based training as a discrete-time dynamical system (dt-DS) and formulates poisoning robustness as a formal safety verification problem. By adapting the concept of barrier certificates (BCs) from control theory, we introduce sufficient conditions to certify a robust radius ensuring that the terminal model remains safe under worst-case ${\ell}_p$-norm-based poisoning. To make this practical, we parameterize BCs as neural networks trained on finite sets of poisoned trajectories. We further derive probably approximately correct (PAC) bounds by solving a scenario convex program (SCP), which yields a confidence lower bound on the certified robustness radius generalizing beyond the training set. Importantly, our framework also extends to certification against test-time attacks, making it the first unified framework to provide formal guarantees in both training and test-time attack settings. Experiments on MNIST, SVHN, CIFAR-10, and CIFAR-100 show that our approach certifies non-trivial perturbation budgets while being model-agnostic and requiring no prior knowledge of the attack or contamination level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。