JKO梯度流在第二阶上隐含减速,提升优化稳定性。
Implicit Bias of the JKO Scheme
- 通过修正能量函数,揭示了JKO方案的二阶隐式偏差
- 在熵、KL散度等常见泛函上,隐式偏差对应物理量如Fisher信息
- 适合研究概率优化中稳定性和收敛性问题的研究者
Wasserstein梯度流为在黎曼流形$(M,g)$的概率测度空间上最小化能量泛函$J$提供了一般框架。其标准时间离散化——Jordan-Kinderlehrer-Otto(JKO)方案,在任意步长$η>0$下生成一列概率分布$ρ_k^η$,以$η$的一阶精度逼近Wasserstein梯度流。但JKO还具备其他一阶积分器不具备的优异性质,例如保持能量耗散,并对$λ$-测地凸泛函$J$具有无条件稳定性。为更深入理解该方案,本文刻画了其在$η$的二阶隐式偏差。我们证明:$ρ_k^η$以$η^2$阶精度逼近修正后的能量$J^η(ρ) = J(ρ) - \fracη{4}\int_M \Big\lVert \nabla_g \frac{δJ}{δρ} (ρ) \Big\rVert_{2}^{2} \,ρ(dx)$上的Wasserstein梯度流,即从$J$中减去$η/4$倍的$J$的平方度量曲率。因此,JKO方案在度量曲率变化剧烈的方向上引入二阶减速。该隐式偏差对应于常见泛函的典型形式:熵对应Fisher信息,KL散度对应Fisher-Hyv{ä}rinen散度,黎曼梯度下降则对应度量$g$下的动能。我们通过若干简单数值例子研究了$J$与$J^η$的差异,包括在Bures-Wasserstein空间上的精确可解朗之万动力学,以及一维四次势能下的朗之万采样。
原文摘要 · Abstract (English)
Wasserstein gradient flow provides a general framework for minimizing an energy functional $J$ over the space of probability measures on a Riemannian manifold $(M,g)$. Its canonical time-discretization, the Jordan-Kinderlehrer-Otto (JKO) scheme, produces for any step size $η>0$ a sequence of probability distributions $ρ_k^η$ that approximate to first order in $η$ Wasserstein gradient flow on $J$. But the JKO scheme also has many other remarkable properties not shared by other first order integrators, e.g. it preserves energy dissipation and exhibits unconditional stability for $λ$-geodesically convex functionals $J$. To better understand the JKO scheme we characterize its implicit bias at second order in $η$. We show that $ρ_k^η$ are approximated to order $η^2$ by Wasserstein gradient flow on a modified energy \[ J^η(ρ) = J(ρ) - \fracη{4}\int_M \Big\lVert \nabla_g \frac{δJ}{δρ} (ρ) \Big\rVert_{2}^{2} \,ρ(dx), \] obtained by subtracting from $J$ the squared metric curvature of $J$ times $η/4$. The JKO scheme therefore adds at second order in $η$ a deceleration in directions where the metric curvature of $J$ is rapidly changing. This corresponds to canonical implicit biases for common functionals: for entropy the implicit bias is the Fisher information, for KL-divergence it is the Fisher-Hyv{ä}rinen divergence, and for Riemannian gradient descent it is the kinetic energy in the metric $g$. To understand the differences between minimizing $J$ and $J^η$ we study JKO-Flow, Wasserstein gradient flow on $J^η$, in several simple numerical examples. These include exactly solvable Langevin dynamics on the Bures-Wasserstein space and Langevin sampling from a quartic potential in 1D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。