将几何与概率结合,用流形投影解释生成模型,提升图像去噪与生成效果。
A Geometric Unification of Generative AI with Manifold-Probabilistic Projection Models
- 引入距离函数统一显式与隐式流形描述,融合几何与概率建模。
- 提出新模型LMPPM,在多个数据集上优于传统扩散模型。
- 为扩散模型提供几何解释,适合研究生成机制与图像重建的学者。
大多数图像生成模型假设图像本质上是高维空间中的低维对象,且主题图像数据集构成光滑或分段光滑的流形。现有方法常忽视几何结构,仅依赖概率方法,通过核方法等通用逼近技术近似分布。部分生成模型虽引入低维潜在空间,但其上的概率分布被视为平凡且预设为均匀分布。本文针对盲图像去噪(BID)问题,提出一种融合几何与概率视角的新框架。该框架改进了传统概率方法,通过引入几何假设使核方法更有效;同时扩展了先前几何方法,利用距离函数结合显式与隐式流形表示。所提框架将扩散模型解释为向‘优质图像’流形的投影过程,由此构建出新的确定性模型——流形-概率投影模型(MPPM),可在像素空间与潜在空间中运行。实验表明,潜在空间版本的LMPPM在多个数据集上均优于潜在扩散模型(LDM),在图像恢复与生成任务中表现更优。
原文摘要 · Abstract (English)
Most models of generative AI for images assume that images are inherently low-dimensional objects embedded within a high-dimensional space. Additionally, it is often implicitly assumed that thematic image datasets form smooth or piecewise smooth manifolds. Common approaches overlook the geometric structure and focus solely on probabilistic methods, approximating the probability distribution through universal approximation techniques such as the kernel method. In some generative models the low dimensional nature of the data manifest itself by the introduction of a lower dimensional latent space. Yet, the probability distribution in the latent or the manifold's coordinate space is considered uninteresting and is predefined or considered uniform. In this study, we address the problem of Blind Image Denoising (BID), and to some extent, the problem of generating images from noise by unifying geometric and probabilistic perspectives. We introduce a novel framework that improves upon existing probabilistic approaches by incorporating geometric assumptions that enable the effective use of kernel-based probabilistic methods. Furthermore, the proposed framework extends prior geometric approaches by combining explicit and implicit manifold descriptions through the introduction of a distance function. The resulting framework demystifies diffusion models by interpreting them as a projection mechanism onto the manifold of ``good images''. This interpretation leads to the construction of a new deterministic model, the Manifold-Probabilistic Projection Model (MPPM), which operates in both the representation (pixel) space and the latent space. We demonstrate that the Latent MPPM (LMPPM) outperforms the Latent Diffusion Model (LDM) across various datasets, achieving superior results in terms of image restoration and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。