arXiv:2502.03795cs.LGmath.CA2025-02被引 3

用神经微分方程学习概率分布,给出精度与网络规模的理论保证。

Distribution learning via neural differential equations: minimal energy regularization and approximation theory

  • 通过时间依赖速度场实现映射的直线插值,最小化能量正则化目标。
  • 速度场光滑性受源/目标分布平滑性控制,误差ε对应网络规模有显式上界。
  • 适用于生成模型、密度估计等任务,适合关注理论保障的研究者。

神经常微分方程(Neural ODEs)能以可逆传输映射的形式表达复杂概率分布,用于生成建模、密度估计和贝叶斯推断。本文证明:对一大类传输映射 $T$,存在时变速度场可实现其诱导位移的直线插值 $(1-t)x + tT(x)$,$t \in [0,1]$。该速度场恰好是最小化特定最小能量正则化训练目标的解。我们推导出速度场 $C^k$ 范数的显式上界,其与 $T$ 的 $C^k$ 范数呈多项式关系;对于三角形(Knothe--Rosenblatt)映射,上界还与源/目标密度的 $C^k$ 范数多项式相关。结合分布逼近的稳定性分析,证明在任意精度 $\varepsilon > 0$ 下,可通过深度神经网络表示的速度场实现对目标分布的 Wasserstein 或 Kullback--Leibler 近似,且网络大小在 $\varepsilon$、维度及源/目标密度光滑性下有明确上界。同一神经网络结构也提供正则化训练目标值的保证。

原文摘要 · Abstract (English)

Neural ordinary differential equations (ODEs) provide expressive representations of invertible transport maps that can be used to approximate complex probability distributions, e.g., for generative modeling, density estimation, and Bayesian inference. We show that for a large class of transport maps $T$, there exists a time-dependent ODE velocity field realizing a straight-line interpolation $(1-t)x + tT(x)$, $t \in [0,1]$, of the displacement induced by the map. Moreover, we show that such velocity fields are minimizers of a training objective containing a specific minimum-energy regularization. We then derive explicit upper bounds for the $C^k$ norm of the velocity field that are polynomial in the $C^k$ norm of the corresponding transport map $T$; in the case of triangular (Knothe--Rosenblatt) maps, we also show that these bounds are polynomial in the $C^k$ norms of the associated source and target densities. Combining these results with stability arguments for distribution approximation via ODEs, we show that Wasserstein or Kullback--Leibler approximation of the target distribution to any desired accuracy $ε> 0$ can be achieved by a deep neural network representation of the velocity field whose size is bounded explicitly in terms of $ε$, the dimension, and the smoothness of the source and target densities. The same neural network ansatz yields guarantees on the value of the regularized training objective.

神经微分方程分布学习理论分析生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。