arXiv:2601.14519cs.LGcs.AI2026-01

探究对抗攻击是否真实反映模型对随机扰动的脆弱性

How Worst-Case Are Adversarial Attacks? Linking Adversarial and Perturbation Robustness

  • 用浓度参数κ统一建模从随机噪声到对抗方向的扰动分布
  • 在ImageNet和CIFAR-10上验证:对抗攻击不总能代表随机扰动风险
  • 为安全评估中对抗攻击的可信度提供实证依据

对抗攻击被广泛用于发现模型漏洞,但其能否作为随机扰动鲁棒性的有效代理仍存争议。本文提出一种概率分析框架,通过浓度参数κ量化不同方向性扰动下的误判风险,该参数在各向同性噪声与对抗方向之间连续插值。基于此,设计了一种旨在探测更接近均匀噪声区域脆弱性的攻击策略。在ImageNet和CIFAR-10上的系统实验对比多种攻击方法,揭示了对抗攻击在何种条件下能有意义地反映对随机扰动的鲁棒性,何种条件下不能,从而为面向安全的鲁棒性评估中对抗攻击的使用提供指导。

原文摘要 · Abstract (English)

Adversarial attacks are widely used to identify model vulnerabilities; however, their validity as proxies for robustness to random perturbations remains debated. We ask whether an adversarial example provides a representative estimate of misprediction risk under stochastic perturbations of the same magnitude, or instead reflects an atypical worst-case event. To address this question, we introduce a probabilistic analysis that quantifies this risk with respect to directionally biased perturbation distributions, parameterized by a concentration factor $κ$ that interpolates between isotropic noise and adversarial directions. Building on this, we study the limits of this connection by proposing an attack strategy designed to probe vulnerabilities in regimes that are statistically closer to uniform noise. Experiments on ImageNet and CIFAR-10 systematically benchmark multiple attacks, revealing when adversarial success meaningfully reflects robustness to perturbations and when it does not, thereby informing their use in safety-oriented robustness evaluation.

对抗攻击鲁棒性图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。