用大模型生成复杂对比图像,提升文生图模型的组合理解能力
Progressive Compositionality in Text-to-Image Generative Models
- 用大模型构造复杂场景,结合VQA自动构建1.5万对对比图像数据集
- 提出多阶段课程学习方法,让扩散模型从难例中有效学习组合关系
- 显著提升复杂组合场景下的文生图质量,适合研究视觉语言对齐的学者
尽管扩散模型在文本到图像生成方面表现优异,但在理解物体与属性之间的组合关系时仍存在困难,尤其在复杂场景下。现有方法通过优化交叉注意力机制或利用语义变化极小的图文对进行训练来缓解问题。然而,能否直接生成高质量、可被扩散模型基于视觉表示区分的对比图像?本文利用大语言模型生成真实且复杂的场景,结合视觉问答系统与扩散模型,自动构建包含1.5万对高质量对比图像的数据集ConPair。这些图像对具有微小的视觉差异,涵盖广泛属性类别,尤其聚焦于复杂自然场景。为有效学习这些错误案例(即难负样本),我们提出EvoGen——一种针对扩散模型的多阶段课程对比学习框架。在多种组合场景上的大量实验表明,该框架在组合性文生图基准测试中具有显著有效性。
原文摘要 · Abstract (English)
Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes, especially in complex settings. Existing solutions have tackled these challenges by optimizing the cross-attention mechanism or learning from the caption pairs with minimal semantic changes. However, can we generate high-quality complex contrastive images that diffusion models can directly discriminate based on visual representations? In this work, we leverage large-language models (LLMs) to compose realistic, complex scenarios and harness Visual-Question Answering (VQA) systems alongside diffusion models to automatically curate a contrastive dataset, ConPair, consisting of 15k pairs of high-quality contrastive images. These pairs feature minimal visual discrepancies and cover a wide range of attribute categories, especially complex and natural scenarios. To learn effectively from these error cases, i.e., hard negative images, we propose EvoGen, a new multi-stage curriculum for contrastive learning of diffusion models. Through extensive experiments across a wide range of compositional scenarios, we showcase the effectiveness of our proposed framework on compositional T2I benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。