根据人眼感知效率动态调整图像生成步数,省时又节能
BudgetFusion: Perceptually-Guided Adaptive Diffusion Models
- 通过预测多层级感知指标,动态决定每张图该用多少生成步数
- 在Stable Diffusion上实测每提示节省5秒,感知质量无损失
- 适合关注生成效率与能耗平衡的研究者和开发者
扩散模型在文本到图像生成任务中表现出前所未有的成功。尽管这些模型能生成高质量、逼真的图像,但其逐步去噪过程带来了高昂的计算开销和能源消耗。为应对这一问题,已有多种方法提升推理效率,但大多采用固定的神经网络简化或文本提示优化策略。我们观察到:不同文本提示生成的图像,所需计算量可能不同,且并非所有去噪步骤对人眼感知都有同等价值。基于此,我们提出BudgetFusion——一种新颖的模型,可预测图像生成前最高效的扩散步数,以实现感知最优。该方法通过预测多层级感知指标与扩散步数的关系来实现。以Stable Diffusion为例,我们进行了数值分析和用户研究。实验表明,BudgetFusion可在不降低感知相似度的前提下,每提示节省最多5秒时间。我们希望本工作能推动思考一个核心问题:生成模型每消耗一瓦特能量,人类感知上能得到多少提升?
原文摘要 · Abstract (English)
Diffusion models have shown unprecedented success in the task of text-to-image generation. While these models are capable of generating high-quality and realistic images, the complexity of sequential denoising has raised societal concerns regarding high computational demands and energy consumption. In response, various efforts have been made to improve inference efficiency. However, most of the existing efforts have taken a fixed approach with neural network simplification or text prompt optimization. Are the quality improvements from all denoising computations equally perceivable to humans? We observed that images from different text prompts may require different computational efforts given the desired content. The observation motivates us to present BudgetFusion, a novel model that suggests the most perceptually efficient number of diffusion steps before a diffusion model starts to generate an image. This is achieved by predicting multi-level perceptual metrics relative to diffusion steps. With the popular Stable Diffusion as an example, we conduct both numerical analyses and user studies. Our experiments show that BudgetFusion saves up to five seconds per prompt without compromising perceptual similarity. We hope this work can initiate efforts toward answering a core question: how much do humans perceptually gain from images created by a generative model, per watt of energy?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。