arXiv:2606.21488cs.LG2026-06

证明了对抗训练无法简化为正则化,深层网络仍难高效防御对抗样本。

Robustness Cannot be Reduced to Regularization: Studying Adversarial Training Beyond the Linear Case

  • 通过归约法证明两层网络不存在对抗风险与正则化风险的等价关系
  • 实验证明在Wide-ResNet上该不可能性依然存在,深度模型难以简化训练
  • 揭示对抗鲁棒性本质不同于传统正则化,适合关注安全性的研究者

机器学习模型对对抗样本的脆弱性已成为重大问题。尽管对抗训练是最有效的应对方法之一,但其高昂的计算成本仍是实际部署的障碍。近期降低计算成本的研究在线性模型中依赖于对抗风险与某种简化正则化风险之间的形式等价性,从而实现更高效的训练。这自然引出一个问题:这种等价性能否推广到非线性模型?本文正式证明:对于两层神经网络,这种等价性不可能存在。证明基于关键性质的归约,这些性质从根本上将对抗风险与仅具弱数据依赖性的简单正则化风险区分开。此外,我们在Wide-ResNets上提供了经验证据,表明此类不可能性在更深、更具表达力的架构中依然成立。

原文摘要 · Abstract (English)

The vulnerability of ML models to adversarial examples has recently emerged as a major concern. While adversarial training is one of the most effective countermeasures to this issue, its high computational cost remains an obstacle to practical deployment. Recent progress in reducing this cost has relied, in the case of linear models, on a formal equivalence between the adversarial risk and a simpler form of regularized risk. This enabled significantly more efficient training procedures, which naturally raises the question of whether such an equivalence can be extended beyond linear models. In this work, we formally show that no such equivalence is possible for two-layer networks. Our proofs proceed via a reduction to key properties that fundamentally separate the adversarial risk from any simple regularized risk which would only exhibit a weak form of data dependence. Beyond this setting, we provide empirical evidence on Wide-ResNets indicating that the same type of impossibility persists in deeper and more expressive architectures.

对抗训练模型鲁棒性神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。