arXiv:2411.05079cs.CVcs.CL2024-11EMNLP

高精度描述比多数量描述更关键,合成数据可替代人工标注。

Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model

  • 用大视觉语言模型生成合成图像描述以提升训练质量
  • 人工标注中精确度对图文对齐影响大于召回率
  • 合成数据训练的模型表现接近真实数据,适合资源有限场景

尽管文本到图像模型取得进展,但生成与文本描述精确匹配的图像仍具挑战性,主要源于训练数据的图文错位。本文分析了图像描述中的精度与召回率在训练中的关键作用。对人工标注描述的分析表明,两者均重要,但精度影响更大。基于此,我们利用大视觉语言模型生成合成描述用于训练。结果显示,使用合成描述训练的模型行为与使用人工标注数据的模型相似,凸显了合成数据在文本到图像训练中的潜力。

原文摘要 · Abstract (English)

Despite advancements in text-to-image models, generating images that precisely align with textual descriptions remains challenging due to misalignment in training data. In this paper, we analyze the critical role of caption precision and recall in text-to-image model training. Our analysis of human-annotated captions shows that both precision and recall are important for text-image alignment, but precision has a more significant impact. Leveraging these insights, we utilize Large Vision Language Models to generate synthetic captions for training. Models trained with these synthetic captions show similar behavior to those trained on human-annotated captions, underscores the potential for synthetic data in text-to-image training.

图文对齐合成数据视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。