arXiv:2503.00240cs.LGcs.AI2025-03被引 2

提出1-Lipschitz网络初始化新问题:深层网络会自动衰减至零。

1-Lipschitz Network Initialization for Certifiably Robust Classification Applications: A Decay Problem

  • 通过分析AOL与SLL架构,发现权重方差不影响输出分布。
  • 证明深度1-Lipschitz网络在标准初始化下必然衰减至零。
  • 适合关注对抗鲁棒性模型稳定性的研究者阅读。

本文研究两种典型1-Lipschitz网络架构——Almost-Orthogonal-Layers (AOL) 和 SDP-based Lipschitz Layers (SLL) 的权重参数化方式,分析其对深层1-Lipschitz前馈网络初始化的影响,并揭示背后的潜在问题。这些网络主要用于可认证鲁棒分类,通过限制扰动对分类输出的影响来抵御对抗攻击。在标准正态分布初始化下,计算了参数化权重方差的精确值与上界;进一步推广到广义正态分布,覆盖均匀、拉普拉斯和正态分布初始化。结果表明,权重方差对输出方差分布无影响,仅权重矩阵维度起作用。此外,本文证明了在任意标准初始化下,深层1-Lipschitz网络总会衰减至零。

原文摘要 · Abstract (English)

This paper discusses the weight parametrization of two standard 1-Lipschitz network architectures, the Almost-Orthogonal-Layers (AOL) and the SDP-based Lipschitz Layers (SLL). It examines their impact on initialization for deep 1-Lipschitz feedforward networks, and discusses underlying issues surrounding this initialization. These networks are mainly used in certifiably robust classification applications to combat adversarial attacks by limiting the impact of perturbations on the classification output. Exact and upper bounds for the parameterized weight variance were calculated assuming a standard Normal distribution initialization; additionally, an upper bound was computed assuming a Generalized Normal Distribution, generalizing the proof for Uniform, Laplace, and Normal distribution weight initializations. It is demonstrated that the weight variance holds no bearing on the output variance distribution and that only the dimension of the weight matrices matters. Additionally, this paper demonstrates that the weight initialization always causes deep 1-Lipschitz networks to decay to zero.

对抗鲁棒网络初始化1-Lipschitz深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。