arXiv:2604.25289cs.LGcs.CV2026-04被引 1

无需时间条件也能生成高质量图像,关键在数据流形的几何对齐。

Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds

论文配图:Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds
图 1 · 摘自论文原文
  • 从几何视角分析扩散过程中的噪声数据流形,发现其集中在低维柱面结构上。
  • 改造DDIM前向过程使其符合流匹配方法,实现无时间条件下的高质量生成。
  • 可扩展至类别条件生成,用类别分离的时间空间替代传统条件嵌入。

实际训练扩散模型通常需要显式的时间条件以指导去噪采样过程。尤其在确定性方法如DDIM中,缺少时间条件会导致性能显著下降。然而,其他确定性采样方法(如流匹配)可在无时间条件的情况下生成高质量内容,引发对其必要性的质疑。本文从几何角度重新审视时间条件的作用:分析前向扩散过程中噪声数据分布的演化,证明在高维空间中这些分布集中于嵌入输入空间的低维柱面状流形。我们认为,成功生成的关键在于高维空间中这些流形的解耦。基于此洞察,我们修改了DDIM的前向过程,使其与流匹配方法对齐,证明只要噪声流形按流匹配方式演化,即使无时间条件也能实现高质量生成。此外,我们通过将类别解耦到独立的时间空间,将该框架扩展至类别条件生成,实现使用类别无关的去噪模型进行类别条件合成。大量实验验证了理论分析,并表明无需显式条件嵌入即可实现高质量生成。

原文摘要 · Abstract (English)

Practically, training diffusion models typically requires explicit time conditioning to guide the network through the denoising sampling process. Especially in deterministic methods like DDIM, the absence of time conditioning leads to significant performance degradation. However, other deterministic sampling approaches, such as flow matching, can generate high-quality content without this conditioning, raising the question of its necessity. In this work, we revisit the role of time conditioning from a geometric perspective. We analyze the evolution of noisy data distributions under the forward diffusion process and demonstrate that, in high-dimensional spaces, these distributions concentrate on low-dimensional hyper-cylinder-like manifolds embedded within the input space. Successful generation, we argue, stems from the disentanglement of these manifolds in high-dimensional space. Based on this insight, we modify the forward process of DDIM to align the noisy data manifold with the flow-matching approach, proving that DDIM can generate high-quality content without time conditioning, provided the noisy manifold evolves according to the flow-matching method. Additionally, we extend our framework to class-conditioned generation by decoupling classes into distinct time spaces, enabling class-conditioned synthesis with a class-unconditional denoising model. Extensive experiments validate our theoretical analysis and show that high-quality generation is achievable without explicit conditional embeddings.

扩散模型流匹配生成建模几何分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。