用流体物理方程建模图像生成,让扩散模型更真实多样。
Beyond Blur: A Fluid Perspective on Generative Diffusion Models
- 基于流体输运方程设计图像退化过程,融合定向流动与扩散。
- 生成图像多样性提升,色彩保持稳定,优于传统方法。
- 适合对物理机制与生成质量有追求的研究者。
我们提出一种基于对流-扩散过程的新型偏微分方程(PDE)驱动图像退化流程,用于生成式图像合成。该前向过程通过耦合方向性对流、各向同性扩散和高斯噪声的物理可解释性PDE实现,受无量纲数(佩克莱特数、傅里叶数)控制。采用GPU加速的自定义格子玻尔兹曼求解器实现数值计算,以快速评估。为引入真实湍流效果,生成具有相干运动和多尺度混合特性的随机速度场。在生成过程中,神经网络学习反转该对流-扩散算子,构成一种新型生成模型。我们证明此前方法可视为该算子的特例,表明本框架能统一现有基于PDE的退化技术。实验显示,对流成分显著提升生成图像的多样性与质量,同时保持整体色域不变。本工作连接流体动力学、无量纲PDE理论与深度生成建模,为基于扩散的合成提供了物理启发的新视角。
原文摘要 · Abstract (English)
We propose a novel PDE-driven corruption process for generative image synthesis based on advection-diffusion processes which generalizes existing PDE-based approaches. Our forward pass formulates image corruption via a physically motivated PDE that couples directional advection with isotropic diffusion and Gaussian noise, controlled by dimensionless numbers (Peclet, Fourier). We implement this PDE numerically through a GPU-accelerated custom Lattice Boltzmann solver for fast evaluation. To induce realistic turbulence, we generate stochastic velocity fields that introduce coherent motion and capture multi-scale mixing. In the generative process, a neural network learns to reverse the advection-diffusion operator thus constituting a novel generative model. We discuss how previous methods emerge as specific cases of our operator, demonstrating that our framework generalizes prior PDE-based corruption techniques. We illustrate how advection improves the diversity and quality of the generated images while keeping the overall color palette unaffected. This work bridges fluid dynamics, dimensionless PDE theory, and deep generative modeling, offering a fresh perspective on physically informed image corruption processes for diffusion-based synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。