证明流模型必须分步生成才能通用建模,单步不行。
Incremental Generation is Necessary and Sufficient for Universality in Flow-Based Modelling
- 用拓扑动力学证明单步流模型无法覆盖所有自然变换
- 分步组合流可实现 $O(n^{-1/d})$ 精度逼近,维度依赖有限
- 适用于需要精确分布生成的高维数据建模任务
增量式流模型重塑了生成建模,但其经验优势缺乏严格的逼近理论基础。本文证明:在 $[0,1]^d$ 上所有与去噪流程兼容的保向同胚映射中,增量生成既是必要也是充分条件。所有保证均对底层映射一致成立,意味着样本和分布层面的统一逼近。我们首次通过新的拓扑-动力学论证证明:无论网络架构、宽度、深度或 Lipschitz 激活函数如何,单步自治流的集合是贫瘠的,因此不具普遍性。相反,利用自治流的代数性质,我们证明每个保向 Lipschitz 同胚均可由至多 $K_d$ 个此类流的复合以 $O(n^{-1/d})$ 的速率逼近,其中 $K_d$ 仅依赖于维度。在额外光滑性假设下,逼近率可摆脱维度依赖,且 $K_d$ 可在被逼近类上统一选取。最后,通过线性升维,我们获得了连续函数和概率测度(作为经验测度的前推)的结构化通用逼近结果,其 $1$-Wasserstein 误差趋于零。
原文摘要 · Abstract (English)
Incremental flow-based denoising models have reshaped generative modelling, but their empirical advantage still lacks a rigorous approximation-theoretic foundation. We show that incremental generation is necessary and sufficient for universal flow-based generation on the largest natural class of self-maps of $[0,1]^d$ compatible with denoising pipelines, namely the orientation-preserving homeomorphisms of $[0,1]^d$. All our guarantees are uniform on the underlying maps and hence imply approximation both samplewise and in distribution. Using a new topological-dynamical argument, we first prove an impossibility theorem: the class of all single-step autonomous flows, independently of the architecture, width, depth, or Lipschitz activation of the underlying neural network, is meagre and therefore not universal in the space of orientation-preserving homeomorphisms of $[0,1]^d$. By exploiting algebraic properties of autonomous flows, we conversely show that every orientation-preserving Lipschitz homeomorphism on $[0,1]^d$ can be approximated at rate $O(n^{-1/d})$ by a composition of at most $K_d$ such flows, where $K_d$ depends only on the dimension. Under additional smoothness assumptions, the approximation rate can be made dimension-free, and $K_d$ can be chosen uniformly over the class being approximated. Finally, by linearly lifting the domain into one higher dimension, we obtain structured universal approximation results for continuous functions and for probability measures on $[0,1]^d$, the latter realized as pushforwards of empirical measures with vanishing $1$-Wasserstein error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。