揭示高维分类中持续对抗攻击的存在性及其增长规律
On the Existence of Consistent Adversarial Attacks in High-Dimensional Linear Classification
- 提出新误差度量,区分对抗攻击与模型表达能力不足
- 理论证明越过参数化的模型对保持标签的扰动越敏感
- 为理解对抗攻击机制提供全新视角,适合安全研究者
什么根本区别了对抗攻击与因模型表达力有限或数据不足导致的误分类?本文在高维二分类设置下研究此问题,其中数据有限带来的统计效应起核心作用。我们引入一种新误差度量,精确刻画这一区别,量化模型对持续对抗攻击(即保持真实标签的扰动)的脆弱性。主要技术贡献是在充分设定模型和隐空间模型下,对这些度量进行严格渐近表征,揭示其与标准鲁棒误差测度不同的脆弱性模式。理论结果表明,随着模型越加过参数化,其对保持标签的扰动的脆弱性上升,为模型对对抗攻击的敏感性提供了理论解释。
原文摘要 · Abstract (English)
What fundamentally distinguishes an adversarial attack from a misclassification due to limited model expressivity or finite data? In this work, we investigate this question in the setting of high-dimensional binary classification, where statistical effects due to limited data availability play a central role. We introduce a new error metric that precisely capture this distinction, quantifying model vulnerability to consistent adversarial attacks -- perturbations that preserve the ground-truth labels. Our main technical contribution is an exact and rigorous asymptotic characterization of these metrics in both well-specified models and latent space models, revealing different vulnerability patterns compared to standard robust error measures. The theoretical results demonstrate that as models become more overparameterized, their vulnerability to label-preserving perturbations grows, offering theoretical insight into the mechanisms underlying model sensitivity to adversarial attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。