优化文本生成图像的训练数据,提升模型对齐与多样性。
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
- 用不同策略生成合成描述文本,测试其对模型影响。
- 随机长度描述能平衡图像质量和文本对齐,不牺牲多样性。
- 描述分布变化会显著改变生成结果的偏向性,适合调优使用。
训练数据是文本到图像模型成功的关键。图像描述的质量和描述性对模型性能至关重要。由于网络爬取数据集存在噪声和不一致问题,近期研究转向使用合成训练描述。尽管这一方法普遍被认为能提升模型能力,但现有文献缺乏对其设计选择的深入分析。本研究系统评估了不同合成描述策略对文本到图像模型下游性能的影响。实验表明,密集且高质量的描述能增强文本对齐,但可能损害输出美观性和多样性;而随机长度的描述在保持样本多样性的同时,实现美学与对齐的平衡提升。此外,描述分布的变化会显著改变模型输出的偏差。研究强调了描述设计在实现最优模型表现中的重要性,并为文本到图像生成的训练数据策略提供了实用指导。
原文摘要 · Abstract (English)
Training data is at the core of any successful text-to-image models. The quality and descriptiveness of image text are crucial to a model's performance. Given the noisiness and inconsistency in web-scraped datasets, recent works shifted towards synthetic training captions. While this setup is generally believed to produce more capable models, current literature does not provide any insights into its design choices. This study closes this gap by systematically investigating how different synthetic captioning strategies impact the downstream performance of text-to-image models. Our experiments demonstrate that dense, high-quality captions enhance text alignment but may introduce trade-offs in output aesthetics and diversity. Conversely, captions of randomized lengths yield balanced improvements across aesthetics and alignment without compromising sample diversity. We also demonstrate that varying caption distributions introduce significant shifts in the output bias of a trained model. Our findings underscore the importance of caption design in achieving optimal model performance and provide practical insights for more effective training data strategies in text-to-image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。