提出新型攻击方法PMA,实现百万级模型鲁棒性评估
Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual Attacks
- 在概率空间定义对抗边界,提升单次攻击效果
- 基于PMA构建高效集成攻击,在百万样本上验证
- 首次完成百万规模白盒鲁棒性评测,揭示评估差距
随着深度学习模型在安全关键场景中的广泛应用,评估其对对抗扰动的脆弱性对于保障可靠性至关重要。过去十年中,大量白盒鲁棒性评估方法(即攻击)被提出,涵盖单步与多步、个体与集成攻击。然而,大规模测试和真实风险反映仍是挑战。本文聚焦图像分类模型,提出一种新型个体攻击方法——概率边际攻击(PMA),在概率空间而非logits空间定义对抗边际。分析表明,PMA优于现有基于交叉熵或logits边际的个体攻击方法。基于PMA,我们设计两种兼顾效率与效果的集成攻击。此外,我们构建了百万规模数据集CC1M(源自CC3M),首次对对抗训练的ImageNet模型进行百万级白盒鲁棒性评估。结果揭示了个体攻击与集成攻击、小规模与百万规模评估之间的显著鲁棒性差距。
原文摘要 · Abstract (English)
As deep learning models are increasingly deployed in safety-critical applications, evaluating their vulnerabilities to adversarial perturbations is essential for ensuring their reliability and trustworthiness. Over the past decade, a large number of white-box adversarial robustness evaluation methods (i.e., attacks) have been proposed, ranging from single-step to multi-step methods and from individual to ensemble methods. Despite these advances, challenges remain in conducting meaningful and comprehensive robustness evaluations, particularly when it comes to large-scale testing and ensuring evaluations reflect real-world adversarial risks. In this work, we focus on image classification models and propose a novel individual attack method, Probability Margin Attack (PMA), which defines the adversarial margin in the probability space rather than the logits space. We analyze the relationship between PMA and existing cross-entropy or logits-margin-based attacks, and show that PMA can outperform the current state-of-the-art individual methods. Building on PMA, we propose two types of ensemble attacks that balance effectiveness and efficiency. Furthermore, we create a million-scale dataset, CC1M, derived from the existing CC3M dataset, and use it to conduct the first million-scale white-box adversarial robustness evaluation of adversarially-trained ImageNet models. Our findings provide valuable insights into the robustness gaps between individual versus ensemble attacks and small-scale versus million-scale evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。