为扩散模型设计无需训练的黎曼度量,让生成路径贴合数据流形。
Be Tangential to Manifold: Discovering Riemannian Metric for Diffusion Models
- 基于得分函数雅可比矩阵的谱结构,构建噪声空间的黎曼度量。
- 生成路径更贴合流形,插值结果更自然,文本图像对齐更好。
- 无需额外训练,适用于图像、视频等多模态生成任务。
扩散模型虽强大,但缺乏显式的低维潜在空间来参数化数据流形,导致难以进行流形感知操作,如几何保真的插值或尊重学习流形的条件引导。本文提出一种无需训练的噪声空间黎曼度量,该度量源自得分函数的雅可比矩阵。关键洞察在于,该雅可比矩阵的谱结构能分离流形的切向与法向方向;我们的度量利用这一分离特性,使生成路径保持切向于流形,避免漂向高密度区域。为验证该度量是否忠实捕捉流形几何,我们从两个互补角度进行检验:其一,该度量下的测地线在合成数据、图像及视频帧数据集上均产生更自然的视觉插值;其二,由该度量诱导的切-法分解可防止无分类器引导过程偏离流形,提升生成质量的同时保持文本-图像对齐性。
原文摘要 · Abstract (English)
Diffusion models are powerful deep generative models, but unlike classical models, they lack an explicit low-dimensional latent space that parameterizes the data manifold. This absence makes it difficult to perform manifold-aware operations, such as geometrically faithful interpolation or conditional guidance that respects the learned manifold. We propose a training-free Riemannian metric on the noise space, derived from the Jacobian of the score function. The key insight is that the spectral structure of this Jacobian separates tangent and normal directions of the data manifold; our metric leverages this separation to encourage paths to stay tangential to the manifold rather than drift toward high-density regions. To validate that our metric faithfully captures the manifold geometry, we examine it from two complementary angles. First, geodesics under our metric yield perceptually more natural interpolations than existing methods on synthetic, image, and video frame datasets. Second, the tangent-normal decomposition induced by our metric prevents classifier-free guidance from deviating off the manifold, improving generation quality while preserving text-image alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。