通过自动调整推理计算分配,显著提升图像生成效率。
Verifier Threshold: An Efficient Test-Time Scaling Approach for Image Generation
- 提出验证器阈值机制,动态优化测试时计算资源分配。
- 在GenEval基准上实现2-4倍计算时间降低,性能不变。
- 适合追求高效图像生成的开发者和部署团队。
图像生成已成为大型生成模型的主流应用。如同测试时计算与推理提升了语言模型能力,图像生成模型也展现出类似优势。特别是针对扩散模型和流模型,通过搜索噪声样本可有效利用测试时计算资源。尽管近期研究尝试在去噪步骤间非均匀分配推理计算预算,但现有方法依赖贪婪启发式策略,常导致计算资源分配低效。本文研究该问题并提出简单改进:验证器阈值(Verifier-Threshold),能自动重分配测试时计算,显著提升效率。在GenEval基准上,相同性能下相较当前最优方法计算时间减少2-4倍。
原文摘要 · Abstract (English)
Image generation has emerged as a mainstream application of large generative models. Just as test-time compute and reasoning have improved language model capabilities, similar benefits have been observed for image generation models. In particular, searching over noise samples for diffusion and flow models has been shown to scale well with test-time compute. While recent works explore allocating non-uniform inference-compute budgets across denoising steps, existing approaches rely on greedy heuristics and often allocate the compute budget ineffectively. In this work, we study this problem and propose a simple fix. We propose Verifier-Threshold, which automatically reallocates test-time compute and delivers substantial efficiency improvements. For the same performance on the GenEval benchmark, we achieve a 2-4x reduction in computational time over the state-of-the-art method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。