arXiv:2507.05300cs.CVcs.AI2025-07被引 1

用结构化标题训练图像生成模型,提升对提示词的遵循度。

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)

  • 用四段式模板统一生成标题:主体、场景、美学、拍摄细节。
  • 在VQA评测中,结构化标题使图文匹配度提升12.7%。
  • 适合需要精准控制生成结果的研究者和开发者。

生成式文本到图像模型常因大规模数据集(如LAION-5B)中标题噪声大、无结构而难以准确理解提示。为此,我们提出在训练中强制使用一致的标题结构,可显著提升模型可控性与对齐能力。本文构建了Re-LAION-Caption 19M,一个高质量子集,包含1900万张1024x1024分辨率图像,其标题由基于Mistral 7B Instruct的LLaVA-Next模型生成,每条标题遵循主体、场景、美学、相机细节四部分模板。我们在PixArt-Σ和Stable Diffusion 2上分别使用结构化与随机打乱的标题进行微调,结果显示结构化版本在视觉问答(VQA)模型上的图文对齐得分更高。该数据集已公开于https://huggingface.co/datasets/supermodelresearch/Re-LAION-Caption19M。

原文摘要 · Abstract (English)

We argue that generative text-to-image models often struggle with prompt adherence due to the noisy and unstructured nature of large-scale datasets like LAION-5B. This forces users to rely heavily on prompt engineering to elicit desirable outputs. In this work, we propose that enforcing a consistent caption structure during training can significantly improve model controllability and alignment. We introduce Re-LAION-Caption 19M, a high-quality subset of Re-LAION-5B, comprising 19 million 1024x1024 images with captions generated by a Mistral 7B Instruct-based LLaVA-Next model. Each caption follows a four-part template: subject, setting, aesthetics, and camera details. We fine-tune PixArt-$Σ$ and Stable Diffusion 2 using both structured and randomly shuffled captions, and show that structured versions consistently yield higher text-image alignment scores using visual question answering (VQA) models. The dataset is publicly available at https://huggingface.co/datasets/supermodelresearch/Re-LAION-Caption19M.

文本生成图像提示词对齐数据增强结构化标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。