arXiv:2510.10000cs.LGmath.OC2025-10

提出更紧的鲁棒性证明和灵活的分布攻击方法,提升神经网络抗干扰能力。

Tight Robustness Certificates and Wasserstein Distributional Attacks for Deep Neural Networks

  • 基于原始问题求解与精确李普希茨证书,改进了广义鲁棒优化上界
  • 在ReLU和现代平滑激活网络中实现可计算的鲁棒性证明
  • 新攻击方法可灵活选择攻击点数量与位置,适用于多种架构

Wasserstein分布鲁棒优化(WDRO)为对抗鲁棒性提供框架,但现有基于全局李普希茨连续性或强对偶性的方法常导致松散上界或计算开销过大。本文采用原始方法,引入精确李普希茨证书以收紧WDRO上界。针对ReLU网络,利用激活单元的分段仿射结构,获得可计算的WDRO问题精确刻画。进一步将分析扩展至具有平滑激活函数(如GELU、SiLU)的现代架构,如Transformer。此外,提出新型Wasserstein分布攻击(WDA, WDA++),可构造最坏情况分布。相比仅限于点扰动的现有攻击,本方法在攻击点数量与位置上更具灵活性。大量实验表明,该框架在对抗主流基线时实现竞争性鲁棒准确率,且证书比现有方法更紧。代码已公开于https://github.com/OLab-Repo/WDA。

原文摘要 · Abstract (English)

Wasserstein distributionally robust optimization (WDRO) provides a framework for adversarial robustness, yet existing methods based on global Lipschitz continuity or strong duality often yield loose upper bounds or require prohibitive computation. We address these limitations with a primal approach and adopt a notion of exact Lipschitz certificates to tighten this upper bound of WDRO. For ReLU networks, we leverage the piecewise-affine structure on activation cells to obtain an exact tractable characterization of the corresponding WDRO problem. We further extend our analysis to modern architectures with smooth activations (e.g., GELU, SiLU), such as Transformers. Additionally, we propose novel Wasserstein Distributional Attacks (WDA, WDA++) that construct candidates for the worst-case distribution. Compared to existing attacks that are restricted to point-wise perturbations, our methods offer greater flexibility in the number and location of attack points. Extensive evaluations demonstrate that our proposed framework achieves competitive robust accuracy against state-of-the-art baselines while offering tighter certificates than existing methods. Our code is available at https://github.com/OLab-Repo/WDA.

鲁棒性分布攻击Wasserstein神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。