arXiv:2505.02255cs.CVcs.AI2025-05被引 1

用合成数据提升蒸馏模型画人脸的逼真度,节省82%算力。

Enhancing AI Face Realism: Cost-Efficient Quality Improvement in Distilled Diffusion Models with a Fully Synthetic Dataset

  • 用合成图像对训练图像翻译头,修复蒸馏模型缺陷。
  • 使FLUX.1-schnell画质逼近FLUX.1-dev,算力降低82%。
  • 适合需要高质人脸生成但资源有限的场景。

本研究提出一种新方法,以提升扩散模型图像生成的性价比。我们假设蒸馏模型(如FLUX.1-schnell)与基线模型(如FLUX.1-dev)之间的差异具有可学习性,尤其在人像生成这类特定领域。为此,我们构建了一个合成配对数据集,并训练了一个快速的图像到图像转换模块。利用两组低质量与高质量的合成图像,模型被训练用于将蒸馏生成器(如FLUX.1-schnell)的输出优化至接近基线模型(如FLUX.1-dev)的视觉质量,后者计算成本更高。实验结果表明,该流水线结合了蒸馏版大模型与增强层,在生成逼真人像方面达到与基线版本相当的效果,同时相比FLUX.1-dev最高可减少82%的计算开销。该研究展示了大规模图像生成中提升AI效率的巨大潜力。

原文摘要 · Abstract (English)

This study presents a novel approach to enhance the cost-to-quality ratio of image generation with diffusion models. We hypothesize that differences between distilled (e.g. FLUX.1-schnell) and baseline (e.g. FLUX.1-dev) models are consistent and, therefore, learnable within a specialized domain, like portrait generation. We generate a synthetic paired dataset and train a fast image-to-image translation head. Using two sets of low- and high-quality synthetic images, our model is trained to refine the output of a distilled generator (e.g., FLUX.1-schnell) to a level comparable to a baseline model like FLUX.1-dev, which is more computationally intensive. Our results show that the pipeline, which combines a distilled version of a large generative model with our enhancement layer, delivers similar photorealistic portraits to the baseline version with up to an 82% decrease in computational cost compared to FLUX.1-dev. This study demonstrates the potential for improving the efficiency of AI solutions involving large-scale image generation.

扩散模型人脸生成模型蒸馏高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。