从平均场理论视角解析对抗训练机制与局限性。
Adversarial Training from Mean Field Perspective
- 基于平均场理论建立新分析框架,无需数据分布假设。
- 推导出不同范数下对抗损失的紧致上界,揭示训练瓶颈。
- 发现无捷径网络难对抗训练,宽度可缓解容量下降问题。
尽管对抗训练在应对对抗样本方面有效,但其训练动态仍不明确。本文首次在无数据分布假设的随机深度神经网络中对对抗训练进行理论分析,提出基于平均场理论的新框架,克服了现有方法的局限性。基于该框架,我们推导出针对不同 $p$、$q$ 值的 $m{ ext{ℓ}_q}$ 范数对抗损失在 $m{ ext{ℓ}_p}$ 范数对抗样本下的经验紧上界。进一步证明:无捷径结构的网络通常无法实现对抗训练,且对抗训练会降低网络容量。同时,我们展示了网络宽度可缓解上述问题,并分析了输入输出维度对上界及权重方差演化的影响。
原文摘要 · Abstract (English)
Although adversarial training is known to be effective against adversarial examples, training dynamics are not well understood. In this study, we present the first theoretical analysis of adversarial training in random deep neural networks without any assumptions on data distributions. We introduce a new theoretical framework based on mean field theory, which addresses the limitations of existing mean field-based approaches. Based on this framework, we derive (empirically tight) upper bounds of $\ell_q$ norm-based adversarial loss with $\ell_p$ norm-based adversarial examples for various values of $p$ and $q$. Moreover, we prove that networks without shortcuts are generally not adversarially trainable and that adversarial training reduces network capacity. We also show that network width alleviates these issues. Furthermore, we present the various impacts of the input and output dimensions on the upper bounds and time evolution of the weight variance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。