让生成模型突破模仿局限,主动创造数据外的新世界。
Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation

- 用谱熵衡量生成分布多样性,设定目标多样性水平
- 低于熵墙时修复模型丢失的多样性,高于熵墙时主动创新
- 无需重训练,可直接用于扩散模型的推理阶段
生成式AI模型通常只模仿数据分布,无法解决生成多样性丧失问题,也无法定义如何超越数据本身的多样性。本文提出想象式生成人工智能(Imaginative Generative AI, IGA),将多样性纳入目标分布设计:在接近参考分布的候选分布中,选择其核协方差算子谱熵达到指定水平的一个。多样性通过固定表示空间中生成分布的核协方差算子的冯诺依曼熵度量,提供一种无参考、基于表示引导的广义概率质量分布方向覆盖程度评估。真实数据分布的谱熵构成“熵墙”。低于熵墙时,IGA执行多样性修复,恢复生成器丢失的变异性,同时保持在数据多样性范围内;高于熵墙时,数据分布不再可行,IGA主动偏离它,生成具有更高表示相对谱多样性的分布,实现操作意义上的“想象生成”。这两个阶段构成从模仿到想象的单一正则化路径,并在每个指定多样性水平上定义一个独立同分布的目标分布。我们建立了熵约束投影的理论,证明在预训练生成器的KL锚定下,最优解满足自洽的指数倾斜关系。该结果导出IGA Guidance——一种无需重训练的推理时方法,适用于基于分数和扩散模型,包括DDPM与DDIM采样器。合成数据与视觉基准测试验证了熵墙以下的多样性修复效果以及墙以上的可控谱外推能力。
原文摘要 · Abstract (English)
Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imaginative Generative AI (IGA), a framework that makes diversity part of the target-distribution design problem: among distributions close to a reference, IGA selects one whose spectral diversity reaches a prescribed level. Diversity is measured by the von Neumann entropy of the generated distribution's kernel covariance operator in a fixed representation space, providing a reference-free representation-guided measure of how broadly probability mass occupies embedding directions. The spectral entropy of the population data distribution defines an Entropy Wall. Below the wall, IGA performs diversity repair, recovering variation that a learned generator has lost while remaining within the diversity level of the data. Beyond the wall, the data distribution itself becomes infeasible, and IGA deliberately departs from it to produce distributions with greater representation-relative spectral diversity, an operational notion of imaginative generation. These regimes form a single regularization path from imitation to imagination and define an i.i.d. target distribution at each prescribed diversity level. We develop the theory of this entropy-constrained projection and show that, under a KL anchor to a pretrained generator, the optimum satisfies a self-consistent exponential-tilt relation. This characterization leads to IGA Guidance, a retraining-free inference-time method for score-based and diffusion models, including DDPM and DDIM samplers. Experiments on synthetic and vision benchmarks demonstrate diversity repair below the Entropy Wall and controlled spectral extrapolation beyond it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。