扩散模型成功的关键是学数据流形,而非完整分布。
When Scores Learn Geometry: Rate Separations under the Manifold Hypothesis
- 从流形几何角度重新解释得分方法,揭示其内在优势。
- 流形学习容错能力比分布学习高σ⁻²量级。
- 适合关注生成模型原理与鲁棒性研究的读者。
基于得分的方法(如扩散模型和贝叶斯逆问题)通常被理解为在低噪声极限(σ→0)下学习数据分布。本文提出新视角:其成功源于隐式学习数据流形,而非完整分布。通过小σ条件下的新分析,我们发现流形信息强度达Θ(σ⁻²),远超分布信息。这表明应从难实现的分布学习转向更可行的几何学习,后者可容忍O(σ⁻²)更大的得分误差。三个推论验证此观点:一、扩散模型仅需o(σ⁻²)的得分误差即可集中于数据支撑集,而恢复具体分布需严格o(1)误差;二、学习流形上均匀分布也更易,可达O(σ⁻²)优势;三、贝叶斯逆问题中最大熵先验比一般先验对得分误差更鲁棒,提升幅度达O(σ⁻²)。初步实验在Stable Diffusion等大模型上验证了理论结论。
原文摘要 · Abstract (English)
Score-based methods, such as diffusion models and Bayesian inverse problems, are often interpreted as learning the data distribution in the low-noise limit ($σ\to 0$). In this work, we propose an alternative perspective: their success arises from implicitly learning the data manifold rather than the full distribution. Our claim is based on a novel analysis of scores in the small-$σ$ regime that reveals a sharp separation of scales: information about the data manifold is $Θ(σ^{-2})$ stronger than information about the distribution. We argue that this insight suggests a paradigm shift from the less practical goal of distributional learning to the more attainable task of geometric learning, which provably tolerates $O(σ^{-2})$ larger errors in score approximation. We illustrate this perspective through three consequences: i) in diffusion models, concentration on data support can be achieved with a score error of $o(σ^{-2})$, whereas recovering the specific data distribution requires a much stricter $o(1)$ error; ii) more surprisingly, learning the uniform distribution on the manifold-an especially structured and useful object-is also $O(σ^{-2})$ easier; and iii) in Bayesian inverse problems, the maximum entropy prior is $O(σ^{-2})$ more robust to score errors than generic priors. Finally, we validate our theoretical findings with preliminary experiments on large-scale models, including Stable Diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。