探究提示复杂度对图像生成质量、多样性与一致性的影响
The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
- 构建新评估框架,系统分析提示复杂度对生成数据的影响
- 复杂提示降低多样性和一致性,但缩小真实与合成数据分布差距
- 提示扩展方法在美学与多样性上超越真实数据,适合高质量生成
文本到图像(T2I)模型可生成无限合成数据,是相比固定有限真实数据的重要资源。以往研究关注合成数据的三大关键属性:质量、多样性与一致性。尽管提示工程是与T2I模型交互的主要方式,但提示复杂度对这些属性的影响尚未系统探索。本文首先通过合成实验揭示提示复杂度泛化难度,并用理论推导解释原因。随后提出新评估框架,对比真实与合成数据的效用,全面分析提示复杂度对主流T2I模型生成数据的影响。研究覆盖CC12M、ImageNet-1k和DCI等多样化数据集,评估不同推理时干预方法。合成实验表明,泛化至更一般条件比反向更难,因扩散模型未学习相关似然估计。大规模实证结果发现,提升提示复杂度导致条件多样性与提示一致性下降,但减少合成与真实数据分布偏移,与合成实验一致。当前推理干预虽能提升生成多样性,却可能使输出超出真实数据支撑范围。其中,提示扩展方法通过预训练语言模型作为似然估计器,持续实现最高图像多样性与美学表现,甚至优于真实数据。
原文摘要 · Abstract (English)
Text-to-image (T2I) models offer great potential for creating virtually limitless synthetic data, a valuable resource compared to fixed and finite real datasets. Previous works evaluate the utility of synthetic data from T2I models on three key desiderata: quality, diversity, and consistency. While prompt engineering is the primary means of interacting with T2I models, the systematic impact of prompt complexity on these critical utility axes remains underexplored. In this paper, we first conduct synthetic experiments to motivate the difficulty of generalization with regard to prompt complexity and explain the observed difficulty with theoretical derivations. Then, we introduce a new evaluation framework that can compare the utility of real data and synthetic data, and present a comprehensive analysis of how prompt complexity influences the utility of synthetic data generated by commonly used T2I models. We conduct our study across diverse datasets, including CC12M, ImageNet-1k, and DCI, and evaluate different inference-time intervention methods. Our synthetic experiments show that generalizing to more general conditions is harder than the other way round, since the former needs an estimated likelihood that is not learned by diffusion models. Our large-scale empirical experiments reveal that increasing prompt complexity results in lower conditional diversity and prompt consistency, while reducing the synthetic-to-real distribution shift, which aligns with the synthetic experiments. Moreover, current inference-time interventions can augment the diversity of the generations at the expense of moving outside the support of real data. Among those interventions, prompt expansion, by deliberately using a pre-trained language model as a likelihood estimator, consistently achieves the highest performance in both image diversity and aesthetics, even higher than that of real data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。