提出新理论框架,揭示图像生成中全局建模易导致记忆而非创造。
Exploring Image Generation via Mutually Exclusive Probability Spaces and Local Correlation Hypothesis
- 构建互斥概率空间与局部相关性假设,从理论上解释生成机制缺陷。
- 实验显示扩大观察范围使自回归模型逐渐转向记忆而非生成。
- 适合关注生成模型本质、对抗过拟合的研究者阅读。
图像生成的主流概率生成模型通常假设学习全局数据分布即可通过采样生成新图像。本文探讨这一核心假设的局限性:学习全局分布反而导致模型记忆而非真正生成。为此提出两个理论框架——互斥概率空间(MEPS)和局部依赖假设(LDH)。MEPS源于随机变量经确定性映射后重叠系数下降的现象,由此推导出重叠系数的下界,并引入二值潜在自编码器(BL-AE),将图像编码为带符号的二进制潜在表示。LDH形式化了有限观测半径内的依赖关系,据此提出γ-自回归随机变量模型(γ-ARVM),该模型以可变观测范围γ预测下一个标记的直方图。实验表明,随着γ增大,自回归模型逐步趋向记忆;在全局依赖极限下,使用BL-AE生成的二值潜在表示时,模型表现如纯记忆体。大量实验与分析支持上述发现。
原文摘要 · Abstract (English)
A common assumption in probabilistic generative models for image generation is that learning the global data distribution suffices to generate novel images via sampling. We investigate the limitation of this core assumption, namely that learning global distributions leads to memorization rather than generative behavior. We propose two theoretical frameworks, the Mutually Exclusive Probability Space (MEPS) and the Local Dependence Hypothesis (LDH), for investigation. MEPS arises from the observation that deterministic mappings (e.g. neural networks) involving random variables tend to reduce overlap coefficients among involved random variables, thereby inducing exclusivity. We further propose a lower bound in terms of the overlap coefficient, and introduce a Binary Latent Autoencoder (BL-AE) that encodes images into signed binary latent representations. LDH formalizes dependence within a finite observation radius, which motivates our $γ$-Autoregressive Random Variable Model ($γ$-ARVM). $γ$-ARVM is an autoregressive model, with a variable observation range $γ$, that predicts a histogram for the next token. Using $γ$-ARVM, we observe that as the observation range increases, autoregressive models progressively shift toward memorization. In the limit of global dependence, the model behaves as a pure memorizer when operating on the binary latents produced by our BL-AE. Comprehensive experiments and discussions support our investigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。