用合成数据提升文本生成图像对物体状态的准确表达
Improving Physical Object State Representation in Text-to-Image Generative Systems
- 自动生成包含多种物体状态的合成数据用于微调
- 在多个基准上平均提升8%以上,特定数据集提升24%+
- 适合关注物理状态建模与真实场景生成的研究者
当前文本到图像生成模型难以准确表现物体状态(如“没有瓶子的桌子”、“空水杯”)。本文设计了一套全自动流程,生成高质量合成数据以精确捕捉物体在不同状态下的特征。随后,我们在该合成数据上微调多个开源文本到图像模型。通过 GPT4o-mini 量化生成图像与提示词的对齐程度,评估显示在公开的 GenAI-Bench 数据集上,四个模型平均绝对提升达 8% 以上。我们还专门构建了 200 个聚焦常见物体在不同物理状态下的提示语,实验表明在该数据集上相比基线平均提升 24% 以上。所有评估提示和代码均已开源。
原文摘要 · Abstract (English)
Current text-to-image generative models struggle to accurately represent object states (e.g., "a table without a bottle," "an empty tumbler"). In this work, we first design a fully-automatic pipeline to generate high-quality synthetic data that accurately captures objects in varied states. Next, we fine-tune several open-source text-to-image models on this synthetic data. We evaluate the performance of the fine-tuned models by quantifying the alignment of the generated images to their prompts using GPT4o-mini, and achieve an average absolute improvement of 8+% across four models on the public GenAI-Bench dataset. We also curate a collection of 200 prompts with a specific focus on common objects in various physical states. We demonstrate a significant improvement of an average of 24+% over the baseline on this dataset. We release all evaluation prompts and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。