用对抗学习动态优化特征空间,提升图像生成质量与分布对齐
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

- 引入可学习的对抗特征空间,动态增强分布对齐能力
- 在多种模型架构和规模下,显著提升生成图像质量与分布匹配度
- 适合追求高保真图像生成的科研与工程人员
弗雷歇距离近年成为生成器后训练中有效的分布级目标,补充了传统的样本级扩散与流匹配损失。然而,直接优化弗雷歇目标可能导致弗雷歇劫持:目标指标持续提升,但视觉质量和其它特征空间中的分布对齐却停滞或下降。我们归因于现有弗雷歇损失使用的静态预训练特征空间,其对真实与生成分布差异的表征不完整且固定。为此,提出对抗式弗雷歇距离(AdvFD),在原始静态弗雷歇目标基础上,加入经校准的对抗学习表示。该表示通过对抗方式最大化真实与生成样本间的弗雷歇差异,而生成器则在由此产生的自适应特征空间中最小化该差异。为防止对抗表示通过特征放大来虚增目标值,进一步引入真实特征白化,规范化其尺度与协方差几何结构,稳定极小极大优化过程。大量实验表明,AdvFD在不同模型规模下,均能一致提升JiT与pMF骨干网络的一步生成器后训练性能。
原文摘要 · Abstract (English)
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving, but visual quality and Fréchet alignment in other feature spaces may stagnate or deteriorate. We attribute this failure to the static pretrained feature spaces used by existing Fréchet losses. These feature spaces provide incomplete and fixed views of the differences between real and generated distributions. To address this limitation, we propose Adversarial Fréchet Distance (AdvFD), which complements the static representation targets in FD-Loss with a calibrated adversarially learned representation. AdvFD augments the original static Fréchet objective with a learnable representation that adversarially maximizes the Fréchet discrepancy between real and generated samples, while the generator minimizes the same discrepancy in the resulting adaptive feature space. To prevent the adversarial representation from trivially increasing the objective through feature amplification, we further introduce real-feature whitening, which normalizes its scale and covariance geometry and stabilizes the min--max optimization. Extensive experiments show that AdvFD consistently improves one-step generator post-training across both JiT and pMF backbones and across different model scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。