让扩散模型根据类别分布生成全新概念,突破传统创意边界。
Distribution-Conditional Generation: From Class Distribution to Creative Generation
- 以类别分布为条件生成图像,实现无语义约束的创造性合成。
- 通过动态概念池与迭代融合,生成符合复杂分布的新型视觉概念。
- 适合需要突破现有语义空间的创意设计、艺术生成场景。
文本到图像(T2I)扩散模型虽能生成语义一致的图像,但受限于训练数据分布,难以合成真正新颖的、分布外的概念。现有方法通常通过组合已知概念提升创意,生成结果虽在分布外,仍受原有语义空间限制。受分类器对模糊输入的软概率输出启发,本文提出分布条件生成(Distribution-Conditional Generation),将创造力建模为基于类别分布的图像合成,实现语义无约束的创意生成。在此基础上,提出DisTok框架:一个编码器-解码器结构,将类别分布映射至隐空间,并解码为创意概念的标记。DisTok维护动态概念池,通过迭代采样与融合概念对,生成与更复杂类别分布对齐的标记。为保证分布一致性,从高斯先验中采样隐向量,解码为标记并生成图像,其类别分布由视觉语言模型预测,用以监督生成标记与视觉语义的一致性。生成标记被加入概念池供后续组合。大量实验表明,DisTok通过统一分布条件融合与基于采样的合成,实现了高效灵活的标记级生成,在文本-图像对齐和人类偏好评分上均达到当前最优表现。
原文摘要 · Abstract (English)
Text-to-image (T2I) diffusion models are effective at producing semantically aligned images, but their reliance on training data distributions limits their ability to synthesize truly novel, out-of-distribution concepts. Existing methods typically enhance creativity by combining pairs of known concepts, yielding compositions that, while out-of-distribution, remain linguistically describable and bounded within the existing semantic space. Inspired by the soft probabilistic outputs of classifiers on ambiguous inputs, we propose Distribution-Conditional Generation, a novel formulation that models creativity as image synthesis conditioned on class distributions, enabling semantically unconstrained creative generation. Building on this, we propose DisTok, an encoder-decoder framework that maps class distributions into a latent space and decodes them into tokens of creative concept. DisTok maintains a dynamic concept pool and iteratively sampling and fusing concept pairs, enabling the generation of tokens aligned with increasingly complex class distributions. To enforce distributional consistency, latent vectors sampled from a Gaussian prior are decoded into tokens and rendered into images, whose class distributions-predicted by a vision-language model-supervise the alignment between input distributions and the visual semantics of generated tokens. The resulting tokens are added to the concept pool for subsequent composition. Extensive experiments demonstrate that DisTok, by unifying distribution-conditioned fusion and sampling-based synthesis, enables efficient and flexible token-level generation, achieving state-of-the-art performance with superior text-image alignment and human preference scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。