给出积分核特征映射的Lipschitz常数显式公式,提升模型稳定性分析能力
Lipschitz bounds for integral kernels
- 基于可微性假设,推导出核特征映射Lipschitz连续的充分条件与显式常数公式
- 证明高斯、ReLU随机神经网络等核的特征映射在特定条件下具有有限Lipschitz常数
- 揭示权重分布二阶矩有限是连续平移不变核Lipschitz性的充要条件,适合理论研究者
核方法与学习理论中,正定核对应的特征映射具有核心作用,其光滑性如Lipschitz连续性直接关联鲁棒性与稳定性。然而,显式的特征映射Lipschitz常数仅在少数情况下可知。本文研究在可微性假设下积分核特征映射的Lipschitz正则性,首先给出保证Lipschitz连续的充分条件并推导出显式常数公式;随后确定特征映射非Lipschitz的条件,并应用于多个重要核类。对于无限宽双层神经网络且权重为各向同性高斯分布的情形,证明其对应核的Lipschitz常数可表示为二维积分上确界,从而对高斯核和ReLU随机神经网络核实现显式刻画。同时研究高斯、拉普拉斯、Matérn等连续平移不变核,其可解释为余弦激活函数的神经网络,证明特征映射Lipschitz连续当且仅当权重分布具有有限二阶矩,并进一步推导其常数。最后提出关于有限宽神经网络中Lipschitz常数收敛渐近行为的开放问题,并通过数值实验验证该现象。
原文摘要 · Abstract (English)
Feature maps associated with positive definite kernels play a central role in kernel methods and learning theory, where regularity properties such as Lipschitz continuity are closely related to robustness and stability guarantees. Despite their importance, explicit characterizations of the Lipschitz constant of kernel feature maps are available only in a limited number of cases. In this paper, we study the Lipschitz regularity of feature maps associated with integral kernels under differentiability assumptions. We first provide sufficient conditions ensuring Lipschitz continuity and derive explicit formulas for the corresponding Lipschitz constants. We then identify a condition under which the feature map fails to be Lipschitz continuous and apply these results to several important classes of kernels. For infinite width two-layer neural network with isotropic Gaussian weight distributions, we show that the Lipschitz constant of the associated kernel can be expressed as the supremum of a two-dimensional integral, leading to an explicit characterization for the Gaussian kernel and the ReLU random neural network kernel. We also study continuous and shift-invariant kernels such as Gaussian, Laplace, and Matérn kernels, which admit an interpretation as neural network with cosine activation function. In this setting, we prove that the feature map is Lipschitz continuous if and only if the weight distribution has a finite second-order moment, and we then derive its Lipschitz constant. Finally, we raise an open question concerning the asymptotic behavior of the convergence of the Lipschitz constant in finite width neural networks. Numerical experiments are provided to support this behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。