arXiv:2409.16541cs.LGmath.AP2024-09被引 2

用索博列夫约束优化低维分布逼近高维概率分布,提升生成模型训练效果。

Monge-Kantorovich Fitting With Sobolev Budgets

  • 以索博列夫范数控制低维映射复杂度,约束支持集结构
  • 提出泛函极小化框架,证明其梯度具有几乎严格单调性
  • 适用于噪声数据流形学习与生成模型正则化分析

当 $m < n$ 时,研究如何用 $m$ 维概率测度 $ν$ 最优逼近 $n$ 维测度 $ρ$,要求 $\mathrm{supp}\ ν$ 的总复杂度有界。当 $ρ$ 集中在 $m$ 维子集附近时,可视为带噪声数据的流形学习问题。通过 $p$-阶 Monge-Kantorovich(Wasserstein)代价 $\mathbb{W}_p^p(ρ, ν)$ 衡量逼近性能,并要求存在映射 $f: \mathbb{R}^m \to \mathbb{R}^n$ 满足 $W^{k,q}$ 索博列夫范数不超过 $\ell \geq 0$,从而将问题转化为在索博列夫预算 $\ell$ 下最小化泛函 $\mathscr J_p(f)$。该问题与 $m=1$ 时的长度约束主曲线及 $k>1$ 时的无监督平滑样条相关但不同。高阶可微性带来新挑战:本文研究 $\mathscr J_p$ 的梯度,定义为称为“重心场”的向量场,并证明其几乎严格单调性。同时提供自然离散化方案并证明其一致性。将其作为生成学习任务的玩具模型,类比提出正则化在改进训练中的新解释。

原文摘要 · Abstract (English)

Given $m < n$, we consider the problem of ``best'' approximating an $n\text{-d}$ probability measure $ρ$ via an $m\text{-d}$ measure $ν$ such that $\mathrm{supp}\ ν$ has bounded total ``complexity.'' When $ρ$ is concentrated near an $m\text{-d}$ set we may interpret this as a manifold learning problem with noisy data. However, we do not restrict our analysis to this case, as the more general formulation has broader applications. We quantify $ν$'s performance in approximating $ρ$ via the Monge-Kantorovich (also called Wasserstein) $p$-cost $\mathbb{W}_p^p(ρ, ν)$, and constrain the complexity by requiring $\mathrm{supp}\ ν$ to be coverable by an $f : \mathbb{R}^{m} \to \mathbb{R}^{n}$ whose $W^{k,q}$ Sobolev norm is bounded by $\ell \geq 0$. This allows us to reformulate the problem as minimizing a functional $\mathscr J_p(f)$ under the Sobolev ``budget'' $\ell$. This problem is closely related to (but distinct from) principal curves with length constraints when $m=1, k = 1$ and an unsupervised analogue of smoothing splines when $k > 1$. New challenges arise from the higher-order differentiability condition. We study the ``gradient'' of $\mathscr J_p$, which is given by a certain vector field that we call the barycenter field, and use it to prove a nontrivial (almost) strict monotonicity result. We also provide a natural discretization scheme and establish its consistency. We use this scheme as a toy model for a generative learning task, and by analogy, propose novel interpretations for the role regularization plays in improving training.

生成模型索博列夫最优传输流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。