对抗样本不是缺陷,而是模型内部表征超叠加的体现。
Adversarial Examples Are Not Bugs, They Are Superposition
- 用超叠加机制解释对抗样本的产生原理。
- 在小模型中操纵超叠加可控制模型鲁棒性。
- 实验证明鲁棒性训练能反向调节超叠加状态。
对抗样本——即经过难以察觉的扰动就可欺骗神经网络的输入——尽管近十年来已有大量研究,仍是深度学习中最令人困惑的现象之一。尽管提出了众多防御方法与解释,但对其根本机制仍无共识。一个未被充分探索的假设是:超叠加(superposition)这一机械可解释性概念,可能是主要成因甚至根本原因。本文提出四条证据支持该假说,显著拓展了Elhage等人(2022)的前期观点:(1) 超叠加可理论解释多种对抗现象;(2) 在小型模型中,干预超叠加可调控模型鲁棒性;(3) 在小型模型中,通过对抗训练调控鲁棒性会反向影响超叠加;(4) 在ResNet18中,对抗训练同样能控制超叠加状态。
原文摘要 · Abstract (English)
Adversarial examples -- inputs with imperceptible perturbations that fool neural networks -- remain one of deep learning's most perplexing phenomena despite nearly a decade of research. While numerous defenses and explanations have been proposed, there is no consensus on the fundamental mechanism. One underexplored hypothesis is that superposition, a concept from mechanistic interpretability, may be a major contributing factor, or even the primary cause. We present four lines of evidence in support of this hypothesis, greatly extending prior arguments by Elhage et al. (2022): (1) superposition can theoretically explain a range of adversarial phenomena, (2) in toy models, intervening on superposition controls robustness, (3) in toy models, intervening on robustness (via adversarial training) controls superposition, and (4) in ResNet18, intervening on robustness (via adversarial training) controls superposition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。