用百个参数优化扩散Transformer,提升图像生成质量与速度
Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration
- 仅用约100个可学习参数,通过进化算法优化扩散Transformer
- 在多类文生图模型上显著提升生成质量,减少推理步数
- 轻量高效,适合资源受限场景下的生成模型部署
本文揭示了扩散Transformer(DiT)在生成任务中的潜在能力。通过对去噪过程的深入分析,我们发现引入单一可学习缩放参数即可显著提升DiT模块性能。基于此洞察,我们提出Calibri,一种参数高效的DiT校准方法。Calibri将校准问题建模为黑箱奖励优化,利用进化算法高效求解,仅修改约100个参数。实验表明,尽管设计轻量,Calibri在多种文生图模型上均持续提升性能,显著减少图像生成所需的推理步数,同时保持高质量输出。
原文摘要 · Abstract (English)
In this paper, we uncover the hidden potential of Diffusion Transformers (DiTs) to significantly enhance generative tasks. Through an in-depth analysis of the denoising process, we demonstrate that introducing a single learned scaling parameter can significantly improve the performance of DiT blocks. Building on this insight, we propose Calibri, a parameter-efficient approach that optimally calibrates DiT components to elevate generative quality. Calibri frames DiT calibration as a black-box reward optimization problem, which is efficiently solved using an evolutionary algorithm and modifies just ~100 parameters. Experimental results reveal that despite its lightweight design, Calibri consistently improves performance across various text-to-image models. Notably, Calibri also reduces the inference steps required for image generation, all while maintaining high-quality outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。