研究核均值差异梯度流的收敛速度,给出精确的数学分析结果。
Quantitative Convergence of Wasserstein Gradient Flows of Kernel Mean Discrepancies
- 基于弱正则性类建立解的存在唯一性,类比二维欧拉方程理论。
- 当s=1时全局指数收敛,s>1时局部多项式收敛,速率依赖s和正则性。
- 首次给出浅层神经网络训练动态的显式局部收敛率,适合理论研究者。
我们研究了核均值差异(KMD,又称最大均值差异,MMD)泛函在Wasserstein梯度流下的定量收敛性。该框架涵盖无限宽浅层神经网络在连续时间极限下的训练动态,以及均场与阻尼极限下具有双粒子瑞兹核相互作用的粒子系统。主要分析针对平方Sobolev距离 $\mathscr{E}^ν_{s}(μ)= \frac{1}{2}\lVert μ-ν\rVert_{\dot H^{-s}}^{2}$ 的情形,其中 $s\geq 1$,$ν$ 为 $d$-维环面上的固定概率测度。首先,受二维欧拉方程杨道维奇理论启发,我们在自然弱正则性类中建立了存在性与唯一性。其次,当 $s=1$ 时,在最小假设下全局指数收敛;当 $s>1$ 时,证明了依赖于 $s$ 及 $μ$、$ν$ 拓扑正则性的局部多项式收敛率,该速率在能量空间与高阶正则性空间中均成立,且对均匀 $ν$ 是紧的。进一步,考虑带ReLU激活的浅层神经网络的群体损失梯度流,其可视为球面 $\mathbb{S}^d$ 上非负测度上的Wasserstein-Fisher-Rao梯度流。通过与 $s=(d+3)/2$ 的Sobolev能量情形对应,导出该动力学的显式多项式局部收敛率。除 $s=1$ 特殊情形外,此前所有这些设置下的非量化收敛性均为开放问题。文中还包含 $d=1$ 维下的数值实验,采用偏微分方程与粒子方法验证分析结果。
原文摘要 · Abstract (English)
We study the quantitative convergence of Wasserstein gradient flows of Kernel Mean Discrepancy (KMD) (also known as Maximum Mean Discrepancy (MMD)) functionals. Our setting covers in particular the training dynamics of shallow neural networks in the infinite-width and continuous time limit, as well as interacting particle systems with pairwise Riesz kernel interaction in the mean-field and overdamped limit. Our main analysis concerns the model case of KMD functionals given by the squared Sobolev distance $ \mathscr{E}^ν_{s}(μ)= \frac{1}{2}\lVert μ-ν\rVert_{\dot H^{-s}}^{2}$ for any $s\geq 1 $ and $ν$ a fixed probability measure on the $d$-dimensional torus. First, inspired by Yudovich theory for the $2d$-Euler equation, we establish existence and uniqueness in natural weak regularity classes. Next, we show that for $s=1$ the flow converges globally at an exponential rate under minimal assumptions, while for $s>1$ we prove local convergence at polynomial rates that depend explicitly on $s$ and on the Sobolev regularity of $μ$ and $ν$. These rates hold both at the energy level and in higher regularity classes and are tight for $ν$ uniform. We then consider the gradient flow of the population loss for shallow neural networks with ReLU activation, which can be cast as a Wasserstein--Fisher--Rao gradient flow on the space of nonnegative measures on the sphere $\mathbb{S}^d$. Exploiting a correspondence with the Sobolev energy case with $s=(d+3)/2$, we derive an explicit polynomial local convergence rate for this dynamics. Except for the special case $s=1$, even non-quantitative convergence was previously open in all these settings. We also include numerical experiments in dimension $d=1$ using both PDE and particle methods which illustrate our analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。