arXiv:2409.10695cs.CVcs.AI2024-09被引 109

用大模型提升图文对齐,生成更精准的视觉设计。

Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

  • 用仅解码器的LLM直接处理文本条件,提升理解能力。
  • 在多个评测中表现领先,能准确还原复杂提示中的文字和颜色。
  • 适合需要精细设计和多语言支持的创意应用开发。

我们提出 Playground v3(PGv3),一种最新的文本到图像生成模型,在多个测试基准上达到当前最优性能,尤其在图形设计方面表现出色并引入新能力。与传统依赖T5或CLIP等预训练文本编码器的模型不同,PGv3完全整合大型语言模型(LLMs),采用新结构,仅使用解码器型LLM提供文本条件。为提升图像描述质量,我们自研了一款可生成多粒度细节描述的标注器,并引入新基准CapsBench评估细粒度图像描述能力。实验表明,PGv3在提示遵循性、复杂推理和文字渲染准确性方面均表现优异。用户偏好研究表明,该模型在贴纸、海报、标志设计等常见设计场景中具备超人类图形设计能力。此外,PGv3还支持精确的RGB色彩控制和强健的多语言理解能力。

原文摘要 · Abstract (English)

We introduce Playground v3 (PGv3), our latest text-to-image model that achieves state-of-the-art (SoTA) performance across multiple testing benchmarks, excels in graphic design abilities and introduces new capabilities. Unlike traditional text-to-image generative models that rely on pre-trained language models like T5 or CLIP text encoders, our approach fully integrates Large Language Models (LLMs) with a novel structure that leverages text conditions exclusively from a decoder-only LLM. Additionally, to enhance image captioning quality-we developed an in-house captioner, capable of generating captions with varying levels of detail, enriching the diversity of text structures. We also introduce a new benchmark CapsBench to evaluate detailed image captioning performance. Experimental results demonstrate that PGv3 excels in text prompt adherence, complex reasoning, and accurate text rendering. User preference studies indicate the super-human graphic design ability of our model for common design applications, such as stickers, posters, and logo designs. Furthermore, PGv3 introduces new capabilities, including precise RGB color control and robust multilingual understanding.

文本生成图像大模型融合图形设计多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。