arXiv:2411.16164cs.CV2024-11综述被引 19

系统梳理十年文本生成图像研究,覆盖440篇论文与关键技术演进

Text-to-Image Synthesis: A Decade Survey

  • 按生成范式分类:从GAN到扩散模型的技术演进路径
  • 总结440项研究,涵盖生成质量、可控性与安全等核心挑战
  • 适合想快速掌握T2I领域全貌的研究者或从业者

当人类阅读特定文本时,常会形成相应的视觉想象,我们期望计算机也能实现这一能力。文本到图像合成(Text-to-Image Synthesis, T2I)旨在从文本描述生成高质量图像,已成为人工智能生成内容(AIGC)的重要方向,也是人工智能研究的变革性议题。基础模型在T2I中发挥关键作用。本文综述了超过440篇近期T2I相关工作。首先简要介绍生成对抗网络(GANs)、自回归模型和扩散模型在图像生成中的应用。在此基础上,讨论这些模型在文本条件下的生成能力与多样性发展。进一步探讨前沿研究在性能、可控性、个性化生成、安全性以及内容与空间关系一致性等方面进展。同时汇总常用数据集与评估指标。最后,展望T2I在AIGC中的潜在应用,并分析当前挑战与未来研究机遇。

原文摘要 · Abstract (English)

When humans read a specific text, they often visualize the corresponding images, and we hope that computers can do the same. Text-to-image synthesis (T2I), which focuses on generating high-quality images from textual descriptions, has become a significant aspect of Artificial Intelligence Generated Content (AIGC) and a transformative direction in artificial intelligence research. Foundation models play a crucial role in T2I. In this survey, we review over 440 recent works on T2I. We start by briefly introducing how GANs, autoregressive models, and diffusion models have been used for image generation. Building on this foundation, we discuss the development of these models for T2I, focusing on their generative capabilities and diversity when conditioned on text. We also explore cutting-edge research on various aspects of T2I, including performance, controllability, personalized generation, safety concerns, and consistency in content and spatial relationships. Furthermore, we summarize the datasets and evaluation metrics commonly used in T2I research. Finally, we discuss the potential applications of T2I within AIGC, along with the challenges and future research opportunities in this field.

文本生成图像AIGC综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。