提出新方法评估模型抗攻击距离,更全面衡量鲁棒性。
A practical approach to evaluating the adversarial distance for machine learning classifiers
- 用迭代攻击与认证结合估算对抗距离上下界。
- 在多个数据集上验证了攻击方法的有效性。
- 适合关注模型安全性的研究人员和工程师。
机器学习分类器的鲁棒性对于真实应用中应对噪声或对抗输入至关重要。准确计算复杂模型和高维数据下的对抗鲁棒性仍具挑战。传统评估通常仅针对特定攻击预算测量对抗精度,信息量有限。本文研究使用迭代对抗攻击与认证方法来估计更具信息量的对抗距离。两者结合可提供对抗鲁棒性的全面评估,计算出对抗距离的上下界。我们展示了可视化结果与消融实验,揭示该方法的应用方式与参数设置。结果显示,所提攻击方法相比现有实现更为有效,而认证方法未达预期。该方法有望推动更有效的机器学习模型鲁棒性评估。
原文摘要 · Abstract (English)
Robustness is critical for machine learning (ML) classifiers to ensure consistent performance in real-world applications where models may encounter corrupted or adversarial inputs. In particular, assessing the robustness of classifiers to adversarial inputs is essential to protect systems from vulnerabilities and thus ensure safety in use. However, methods to accurately compute adversarial robustness have been challenging for complex ML models and high-dimensional data. Furthermore, evaluations typically measure adversarial accuracy on specific attack budgets, limiting the informative value of the resulting metrics. This paper investigates the estimation of the more informative adversarial distance using iterative adversarial attacks and a certification approach. Combined, the methods provide a comprehensive evaluation of adversarial robustness by computing estimates for the upper and lower bounds of the adversarial distance. We present visualisations and ablation studies that provide insights into how this evaluation method should be applied and parameterised. We find that our adversarial attack approach is effective compared to related implementations, while the certification method falls short of expectations. The approach in this paper should encourage a more informative way of evaluating the adversarial robustness of ML classifiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。