arXiv:2410.06895cs.LG2024-10ICML被引 3

ACR指标误导了随机平滑研究,实测发现它无法反映真实鲁棒性。

Average Certified Radius is a Poor Metric for Randomized Smoothing

  • 证明平凡分类器也能获得极高的平均认证半径
  • 训练策略提升ACR时反而降低模型对难样本的鲁棒性
  • 建议改用噪声数据上的基础模型准确率分布作为新指标

随机平滑(RS)是提供对抗攻击认证鲁棒性保障的主流方法。平均认证半径(ACR)被广泛用于追踪进展。本文首次揭示ACR是一个劣质评估指标:理论上证明,一个平凡分类器可拥有任意大的ACR,且ACR对易样本改进极度敏感;此外,基于ACR的比较严重依赖认证预算。实验表明,现有训练策略虽提升ACR,却一致削弱模型在难样本上的鲁棒性。我们提出丢弃难样本、按近似认证半径重加权数据集、极端优化易样本等策略,在不训练全分布鲁棒性的情况下,仍能达成CIFAR-10上最先进的ACR。结果表明ACR已引入强烈偏差,应停止使用。最后建议采用$ p_A $(基础模型在噪声数据上的准确率)的实证分布作为替代指标。

原文摘要 · Abstract (English)

Randomized smoothing (RS) is popular for providing certified robustness guarantees against adversarial attacks. The average certified radius (ACR) has emerged as a widely used metric for tracking progress in RS. However, in this work, for the first time we show that ACR is a poor metric for evaluating robustness guarantees provided by RS. We theoretically prove not only that a trivial classifier can have arbitrarily large ACR, but also that ACR is extremely sensitive to improvements on easy samples. In addition, the comparison using ACR has a strong dependence on the certification budget. Empirically, we confirm that existing training strategies, though improving ACR, reduce the model's robustness on hard samples consistently. To strengthen our findings, we propose strategies, including explicitly discarding hard samples, reweighing the dataset with approximate certified radius, and extreme optimization for easy samples, to replicate the progress in RS training and even achieve the state-of-the-art ACR on CIFAR-10, without training for robustness on the full data distribution. Overall, our results suggest that ACR has introduced a strong undesired bias to the field, and its application should be discontinued in RS. Finally, we suggest using the empirical distribution of $p_A$, the accuracy of the base model on noisy data, as an alternative metric for RS.

随机平滑鲁棒性评估认证半径对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。