arXiv:2507.03450cs.CRcs.AI2025-07被引 1

提出标准化攻击评测框架,解决对抗鲁棒性测试不靠谱问题

Evaluating the Evaluators: Trust in Adversarial Robustness Tests

  • 构建统一测试环境的AttackBench框架,确保攻击评估可复现
  • 通过新优化度量标准,量化比较不同梯度攻击的有效性
  • 适合关注模型鲁棒性验证的研究者与实践者使用

尽管在设计强大的对抗性规避攻击以验证模型鲁棒性方面取得了显著进展,但这些方法的评估常常存在不一致和不可靠的问题。许多评估依赖于不匹配的模型、未经验证的实现以及不均衡的计算预算,可能导致偏差结果并产生虚假的安全感。因此,基于此类有缺陷测试协议的鲁棒性声明可能具有误导性。为提升评估可靠性,本文提出AttackBench,一个用于在标准化、可复现条件下评估基于梯度攻击有效性的基准框架。该框架通过新型最优性度量对现有攻击实现进行排名,帮助研究人员和从业者识别最可靠、高效的攻击方法,用于后续的鲁棒性评估。框架强制执行一致的测试条件,并支持持续更新,成为鲁棒性验证的可靠基础。

原文摘要 · Abstract (English)

Despite significant progress in designing powerful adversarial evasion attacks for robustness verification, the evaluation of these methods often remains inconsistent and unreliable. Many assessments rely on mismatched models, unverified implementations, and uneven computational budgets, which can lead to biased results and a false sense of security. Consequently, robustness claims built on such flawed testing protocols may be misleading and give a false sense of security. As a concrete step toward improving evaluation reliability, we present AttackBench, a benchmark framework developed to assess the effectiveness of gradient-based attacks under standardized and reproducible conditions. AttackBench serves as an evaluation tool that ranks existing attack implementations based on a novel optimality metric, which enables researchers and practitioners to identify the most reliable and effective attack for use in subsequent robustness evaluations. The framework enforces consistent testing conditions and enables continuous updates, making it a reliable foundation for robustness verification.

对抗攻击模型评测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。