arXiv:2607.19855cs.LGcs.CR2026-07

提出最小范数攻击集合,实现可控制查询成本的鲁棒性曲线评估。

Adversarial Frontiers: Minimum-Norm Attack Ensembles for Robustness Evaluation

  • 构建跨范数的最小范数攻击池,系统生成攻击前沿。
  • 在CIFAR-10和ImageNet上,性能优于或等同于AutoAttack。
  • 支持按查询预算动态调整强度,适合模型对比与防御评估。

对抗鲁棒性通常在单一扰动预算ε下,使用预定义攻击集合(如AutoAttack)并限定扰动范数进行评估。但该方法存在根本缺陷:不同模型的鲁棒性-扰动曲线可能交叉或衰减速率不同,导致单ε排名不稳定;现有集合无法证明最优性,存在未知的最坏情况差距;固定攻击配置也无法系统控制攻击强度与评估成本间的权衡。为此,本文提出统一评估框架,基于ℓ₀、ℓ₁、ℓ₂和ℓ∞范数下的最小范数攻击池,构建鲁棒性-扰动曲线。定义攻击前沿为攻击池对模型产生的最坏情况鲁棒性估计。将评估形式化为前沿逼近问题,构造最小范数攻击集合——即从完整池中优化选取的子集,在可控查询预算下逼近前沿,且预算越大,估计越紧。进一步定义防御前沿为模型集在各扰动规模下的最大鲁棒性。提出防御最优性指数,通过衡量防御与防御前沿的差距对防御进行排序,无需指定参考ε。在CIFAR-10和ImageNet上,我们的集合在各预算层级均匹配或超越AutoAttack,且保持固定且可控的查询成本,为从业者提供一种可查询控制、基于曲线的替代方案。

原文摘要 · Abstract (English)

Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single perturbation budget $\varepsilon$ and on a selective choice of perturbation norms. We argue this formulation is fundamentally limited. First, robustness--perturbation curves may intersect or decay at different rates across models, making single-$\varepsilon$ rankings unstable. Second, current ensembles provide no evidence of optimality, leaving an unknown gap to worst-case performance. Third, fixed attack configurations provide no systematic control over the trade-off between attack strength and evaluation cost. To address these limitations, we introduce a unified evaluation framework based on a comprehensive pool of minimum-norm attacks and robustness--perturbation curves across $\ell_0$, $\ell_1$, $\ell_2$ and $\ell_\infty$ norms. We define the attack frontier as the worst-case robustness estimate the attack pool produces against a model. We then formalize evaluation as a frontier-approximation problem, constructing minimum-norm attack ensembles, optimized subsets of the comprehensive pool, that approach the frontier under a controllable query budget, with larger budgets monotonically tightening the estimate. Furthermore, we define the defense frontier as the maximum robustness across the model set at each perturbation size. We finally propose the Defense Optimality Index to rank defenses by their gap to the defense frontier, providing a ranking without selecting a reference $\varepsilon$. On CIFAR-10 and ImageNet, our ensembles match or exceed AutoAttack on most defenses at every budget tier, at fixed and controllable query cost, offering practitioners a query-controlled, curve-based alternative to fixed-$\varepsilon$ evaluation.

对抗攻击鲁棒性评估最小范数防御前沿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。