揭示无噪声条件扩散模型的几何本质,解释其为何稳定生成。
The Geometry of Noise: Why Diffusion Models Don't Need Noise Conditioning
- 提出边际能量概念,构建无噪声条件模型的几何基础
- 证明速度参数化具内在稳定性,可吸收不确定性
- 解析盲模型崩溃原因,给出采样结构稳定性条件
自主生成模型(如平衡匹配与盲扩散)不依赖显式噪声水平条件,学习单一时不变向量场。尽管高维集中使模型能隐式估计噪声水平,但其优化目标与稳定性仍存根本矛盾:当噪声水平视为随机变量时,如何在数据流形附近保持稳定?本文定义边际能量 $E_{\text{marg}}(\mathbf{u}) = -\log p(\mathbf{u})$,其中 $p(\mathbf{u}) = \int p(\mathbf{u}|t)p(t)dt$ 为对未知噪声水平先验积分后的边缘密度。证明自主模型生成实为在该边际能量上的黎曼梯度流。通过新颖的相对能量分解,发现原始边际能量在数据流形法向上存在 $1/t^p$ 奇点,但学习到的时不变场隐含局部共形度量,恰好抵消几何奇点,将无限深势阱转为稳定吸引子。同时建立采样结构稳定性条件。识别出噪声预测参数化的“Jensen间隙”,其作为高增益放大器导致估计误差剧增,解释了确定性盲模型的灾难性失败;反之,速度参数化满足有界增益条件,可将后验不确定性融入平滑几何漂移中,具有固有稳定性。
原文摘要 · Abstract (English)
Autonomous (noise-agnostic) generative models, such as Equilibrium Matching and blind diffusion, challenge the standard paradigm by learning a single, time-invariant vector field that operates without explicit noise-level conditioning. While recent work suggests that high-dimensional concentration allows these models to implicitly estimate noise levels from corrupted observations, a fundamental paradox remains: what is the underlying landscape being optimized when the noise level is treated as a random variable, and how can a bounded, noise-agnostic network remain stable near the data manifold where gradients typically diverge? We resolve this paradox by formalizing Marginal Energy, $E_{\text{marg}}(\mathbf{u}) = -\log p(\mathbf{u})$, where $p(\mathbf{u}) = \int p(\mathbf{u}|t)p(t)dt$ is the marginal density of the noisy data integrated over a prior distribution of unknown noise levels. We prove that generation using autonomous models is not merely blind denoising, but a specific form of Riemannian gradient flow on this Marginal Energy. Through a novel relative energy decomposition, we demonstrate that while the raw Marginal Energy landscape possesses a $1/t^p$ singularity normal to the data manifold, the learned time-invariant field implicitly incorporates a local conformal metric that perfectly counteracts the geometric singularity, converting an infinitely deep potential well into a stable attractor. We also establish the structural stability conditions for sampling with autonomous models. We identify a ``Jensen Gap'' in noise-prediction parameterizations that acts as a high-gain amplifier for estimation errors, explaining the catastrophic failure observed in deterministic blind models. Conversely, we prove that velocity-based parameterizations are inherently stable because they satisfy a bounded-gain condition that absorbs posterior uncertainty into a smooth geometric drift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。