arXiv:2410.08608cs.CVcs.AI2024-10被引 2

对比五种GAN文本生成图像方法,找最优模型。

Text-To-Image with Generative Adversarial Networks

  • 用GAN构建五种文本到图像生成模型
  • 最高生成256×256分辨率图像,最低64×64
  • 通过指标对比选出最优模型,适合图像生成研究者

从人类描述生成真实图像是计算机视觉领域最具挑战性的问题之一。现有文本到图像方法可大致反映描述含义。本文主要目的是对基于生成对抗网络(GAN)的五种不同方法进行简要对比,以实现从文本生成图像。此外,各模型生成图像的分辨率各异,最高可达256×256,最低为64×64。我们还评估并比较了各模型的多项性能指标,通过分析关键指标,确定了该问题下的最优模型。

原文摘要 · Abstract (English)

Generating realistic images from human texts is one of the most challenging problems in the field of computer vision (CV). The meaning of descriptions given can be roughly reflected by existing text-to-image approaches. In this paper, our main purpose is to propose a brief comparison between five different methods base on the Generative Adversarial Networks (GAN) to make image from the text. In addition, each model architectures synthesis images with different resolution. Furthermore, the best and worst obtained resolutions is 64*64, 256*256 respectively. However, we checked and compared some metrics that introduce the accuracy of each model. Also, by doing this study, we found out the best model for this problem by comparing these different approaches essential metrics.

文本生成GAN图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。