arXiv:2505.02527cs.CV2025-05综述被引 13

全面梳理2021-2024年141篇文本生成图像研究,解析主流模型与技术。

Text to Image Generation and Editing: A Survey

  • 按自回归、非自回归、GAN、扩散模型分类梳理基础架构
  • 对比141项研究在数据集、评估指标、训练资源上的表现差异
  • 涵盖生成与编辑两大方向,适合作为领域入门指南

文本到图像生成(T2I)指在文本引导下生成高质量图像。近年来该领域受到广泛关注,涌现出大量研究成果。本文系统回顾了2021至2024年间发表的141篇相关工作。首先介绍T2I的四种基础模型架构(自回归、非自回归、GAN、扩散模型)及常用关键技术(自编码器、注意力机制、无分类器引导)。其次,从生成与编辑两个方向,系统比较各研究使用的编码器与核心技术。同时,基于数据集、评估指标、训练资源与推理速度对研究性能进行横向对比。除四大基础模型外,还涵盖能量模型、近期Mamba架构及多模态方法。此外,探讨了T2I可能带来的社会影响并提出应对方案。最后,总结提升模型性能的关键洞见,并展望未来发展方向。本综述是首个系统性、全面性的T2I综述,旨在为后续研究提供参考,推动该领域持续进步。

原文摘要 · Abstract (English)

Text-to-image generation (T2I) refers to the text-guided generation of high-quality images. In the past few years, T2I has attracted widespread attention and numerous works have emerged. In this survey, we comprehensively review 141 works conducted from 2021 to 2024. First, we introduce four foundation model architectures of T2I (autoregression, non-autoregression, GAN and diffusion) and the commonly used key technologies (autoencoder, attention and classifier-free guidance). Secondly, we systematically compare the methods of these studies in two directions, T2I generation and T2I editing, including the encoders and the key technologies they use. In addition, we also compare the performance of these researches side by side in terms of datasets, evaluation metrics, training resources, and inference speed. In addition to the four foundation models, we survey other works on T2I, such as energy-based models and recent Mamba and multimodality. We also investigate the potential social impact of T2I and provide some solutions. Finally, we propose unique insights of improving the performance of T2I models and possible future development directions. In summary, this survey is the first systematic and comprehensive overview of T2I, aiming to provide a valuable guide for future researchers and stimulate continued progress in this field.

文本生成图像生成综述扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。