arXiv:2410.14591math.FAcs.LG2024-10被引 3

从光滑空间视角重新理解无限宽浅层网络的数学机制

A Lipschitz spaces view of infinitely wide shallow neural networks

  • 用带符号测度和Lipschitz函数对偶刻画网络参数空间
  • 证明在总变差与矩约束下存在最小化解,否则出现异常行为
  • 融合最优传输与再生核空间优势,适用于模型压缩与融合

我们重新审视浅层神经网络的均场参数化,采用定义在无界参数空间上的带符号测度,结合激活函数的正则性与增长性进行对偶配对。该设定直接引出基于Lipschitz函数对偶的非平衡Kantorovich-Rubinstein范数,以及与具有可控增长性的连续函数对偶的测度空间。这些工具揭示了在变分公式中获得最小化解所需的总变差与矩界或惩罚项,并在此条件下证明了强Kantorovich-Rubinstein范数下的紧性结果;在无此条件时,我们展示了多个展示不良行为的例子。此外,Kantorovich-Rubinstein框架使我们能够结合完全线性参数化带来的再生核巴拿赫空间优势与最优传输洞见。我们通过表示定理和经验风险最小化的统一大样本极限,展示了这一协同效应,并提出了用于知识蒸馏与融合的新公式。

原文摘要 · Abstract (English)

We revisit the mean field parametrization of shallow neural networks, using signed measures on unbounded parameter spaces and duality pairings that take into account the regularity and growth of activation functions. This setting directly leads to the use of unbalanced Kantorovich-Rubinstein norms defined by duality with Lipschitz functions, and of spaces of measures dual to those of continuous functions with controlled growth. These allow to make transparent the need for total variation and moment bounds or penalization to obtain existence of minimizers of variational formulations, under which we prove a compactness result in strong Kantorovich-Rubinstein norm, and in the absence of which we show several examples demonstrating undesirable behavior. Further, the Kantorovich-Rubinstein setting enables us to combine the advantages of a completely linear parametrization and ensuing reproducing kernel Banach space framework with optimal transport insights. We showcase this synergy with representer theorems and uniform large data limits for empirical risk minimization, and in proposed formulations for distillation and fusion applications.

神经网络理论最优传输再生核空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。