arXiv:2412.20987cs.LG2024-12被引 2

挑战强防御模型的黑盒攻击,发现现有方法效果有限。

RobustBlack: Challenging Black-Box Adversarial Attacks on State-of-the-Art Defenses

  • 构建新评估框架,测试黑盒攻击对ImageNet上先进防御模型的效果。
  • 顶尖黑盒攻击在简单对抗训练模型前也难奏效,成功率不足20%。
  • 目标模型与代理模型间鲁棒性对齐程度决定迁移攻击成败。

尽管白盒对抗鲁棒性研究已十分深入,但近期的黑盒攻击(包括迁移和查询类方法)主要针对弱防御机制进行评估,未能充分检验其对更先进、中等强度鲁棒模型(如Robustbench榜单上的模型)的有效性。本文质疑黑盒攻击对强防御模型关注不足的问题,提出一个评估框架,用于分析最新黑盒攻击在ImageNet数据集上对顶尖及标准防御机制的攻击效果。实验结果表明:(1)最先进的黑盒攻击即使在面对简单的对抗训练模型时也难以成功;(2)针对强白盒攻击(如AutoAttack)优化的鲁棒模型,同样表现出更强的抗黑盒攻击能力;(3)代理模型与目标模型之间的鲁棒性对齐程度是决定迁移攻击成功率的关键因素。

原文摘要 · Abstract (English)

Although adversarial robustness has been extensively studied in white-box settings, recent advances in black-box attacks (including transfer- and query-based approaches) are primarily benchmarked against weak defenses, leaving a significant gap in the evaluation of their effectiveness against more recent and moderate robust models (e.g., those featured in the Robustbench leaderboard). In this paper, we question this lack of attention from black-box attacks to robust models. We establish a framework to evaluate the effectiveness of recent black-box attacks against both top-performing and standard defense mechanisms, on the ImageNet dataset. Our empirical evaluation reveals the following key findings: (1) the most advanced black-box attacks struggle to succeed even against simple adversarially trained models; (2) robust models that are optimized to withstand strong white-box attacks, such as AutoAttack, also exhibits enhanced resilience against black-box attacks; and (3) robustness alignment between the surrogate models and the target model plays a key factor in the success rate of transfer-based attacks

黑盒攻击对抗样本鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。