用CLIP辅助的GAN实现高效高多样性文本生成图像
Efficiency without Compromise: CLIP-aided Text-to-Image GANs with Increased Diversity
- 设计双判别器结构,结合切片对抗网络提升生成多样性
- 零样本FID媲美大型GAN,训练成本降低百倍
- 提出新评估指标PPD,量化分析每提示下的生成多样性
近期,生成对抗网络(GAN)已成功扩展至十亿规模的图文数据集。然而,大规模训练带来高昂成本,限制了应用与研究。为降低开销,现有方法引入预训练模型,显著减少训练成本,但导致给定提示下的生成多样性大幅下降。为此,本文提出SCAD模型,采用适配图文任务的切片对抗网络(SAN),配备两个专用判别器,在保持生成质量的同时大幅提升多样性。我们还提出“每提示多样性”(Per-Prompt Diversity, PPD)指标,实现对生成多样性的量化评估。SCAD在零样本下达到与最新大型GAN相当的FID表现,训练成本仅为后者的百分之一。
原文摘要 · Abstract (English)
Recently, Generative Adversarial Networks (GANs) have been successfully scaled to billion-scale large text-to-image datasets. However, training such models entails a high training cost, limiting some applications and research usage. To reduce the cost, one promising direction is the incorporation of pre-trained models. The existing method of utilizing pre-trained models for a generator significantly reduced the training cost compared with the other large-scale GANs, but we found the model loses the diversity of generation for a given prompt by a large margin. To build an efficient and high-fidelity text-to-image GAN without compromise, we propose to use two specialized discriminators with Slicing Adversarial Networks (SANs) adapted for text-to-image tasks. Our proposed model, called SCAD, shows a notable enhancement in diversity for a given prompt with better sample fidelity. We also propose to use a metric called Per-Prompt Diversity (PPD) to evaluate the diversity of text-to-image models quantitatively. SCAD achieved a zero-shot FID competitive with the latest large-scale GANs at two orders of magnitude less training cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。