用对抗学习改进流模型,生成图像更逼真。
Continuous Adversarial Flow Models

- 引入可学习判别器替代固定损失,优化训练目标。
- 在ImageNet上使无引导FID从8.26降至3.63,提升显著。
- 适用于已有流模型的后处理,也支持从零训练。
我们提出连续对抗流模型,一种以对抗目标训练的连续时间流模型。与使用固定均方误差准则的流匹配不同,该方法引入可学习判别器引导训练,改变泛化分布,实证表明生成样本更贴近目标数据分布。本方法主要用于对现有流匹配模型进行后训练,亦可从头训练。在ImageNet 256px生成任务中,后训练使潜空间SiT的无引导FID从8.26降至3.63,像素空间JiT从7.17降至3.57;同时改善有引导生成,使SiT的FID从2.06降至1.53,JiT从1.86降至1.80。进一步在文本到图像生成任务中评估,于GenEval和DPG基准上均取得更优结果。
原文摘要 · Abstract (English)
We propose continuous adversarial flow models, a type of continuous-time flow model trained with an adversarial objective. Unlike flow matching, which uses a fixed mean-squared-error criterion, our approach introduces a learned discriminator to guide training. This change in objective induces a different generalized distribution, which empirically produces samples that are better aligned with the target data distribution. Our method is primarily proposed for post-training existing flow-matching models, although it can also train models from scratch. On the ImageNet 256px generation task, our post-training substantially improves the guidance-free FID of latent-space SiT from 8.26 to 3.63 and of pixel-space JiT from 7.17 to 3.57. It also improves guided generation, reducing FID from 2.06 to 1.53 for SiT and from 1.86 to 1.80 for JiT. We further evaluate our approach on text-to-image generation, where it achieves improved results on both the GenEval and DPG benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。