不依赖外部模型,用推理时优化初始噪声提升文生图效果
Performance Plateaus in Inference-Time Scaling for Text-to-Image Diffusion Without External Models
- 通过无外部模型的Best-of-N方法优化扩散模型初始噪声
- 少量优化步骤即可达到算法上限性能,快速达性能平台期
- 适合资源受限场景,如小显存GPU上的高效图像生成
最近研究发现,投入计算资源搜索文本到图像扩散模型的优质初始噪声可提升性能。然而,以往方法需依赖外部模型评估生成图像,无法在显存较小的GPU上运行。为此,我们在多个数据集和骨干网络上,将无需外部模型的Best-of-N推理时缩放方法应用于优化扩散模型初始噪声。结果表明,在此设置下,文生图扩散模型的推理时缩放迅速达到性能平台期,每个算法仅需相对较少的优化步骤即可实现最大可达成性能。
原文摘要 · Abstract (English)
Recently, it has been shown that investing computing resources in searching for good initial noise for a text-to-image diffusion model helps improve performance. However, previous studies required external models to evaluate the resulting images, which is impossible on GPUs with small VRAM. For these reasons, we apply Best-of-N inference-time scaling to algorithms that optimize the initial noise of a diffusion model without external models across multiple datasets and backbones. We demonstrate that inference-time scaling for text-to-image diffusion models in this setting quickly reaches a performance plateau, and a relatively small number of optimization steps suffices to achieve the maximum achievable performance with each algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。