揭秘DiT中adaLN-Zero为何更优,提出改进初始化与新机制
Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

- 发现零初始化是adaLN-Zero性能提升的关键
- 提出adaLN-Gaussian使优化效率稳定提升
- SE-adaLN-Zero在图像与文本生成中均表现更好
扩散变压器(DiT)作为图像生成的新兴架构备受关注,但对其理解仍较浅显。本文深入研究了DiT中关键的条件化机制adaLN-Zero,其性能优于adaLN。我们分析了三种潜在驱动因素:类似SE的结构、零初始化和渐进更新顺序,结果表明零初始化影响最大。基于此,提出分析引导的初始化策略adaLN-Gaussian,既验证了分析,又提升了优化效率。同时,受SE结构启发,提出改进的条件化机制SE-adaLN-Zero。在四个数据集上的大量实验,尤其在ImageNet1K上,证明了两种方法的有效性与泛化能力。此外,在文本到图像生成任务中也验证了其通用性。
原文摘要 · Abstract (English)
Diffusion transformer (DiT), a rapidly emerging architecture for image generation, has gained much attention. However, despite ongoing efforts to improve its performance, the understanding of DiT remains superficial. In this work, we delve into and investigate a critical conditioning mechanism within DiT, adaLN-Zero, which achieves superior performance compared to adaLN. Our work studies three potential elements driving this performance, including an SE-like structure, zero-initialization, and a "gradual" update order, among which zero-initialization is proved to be the most influential. Building on this understanding, we propose an analysis-guided initialization strategy, termed adaLN-Gaussian, which serves both as an empirical validation of our analysis and as a practical initialization method that consistently improves optimization efficiency. On the other hand, inspired by the SE-like structure, we introduce an improved conditioning mechanism called SE-adaLN-Zero. Extensive experiments following DiT on four datasets, especially on ImageNet1K demonstrate the effectiveness and generalization of adaLN-Gaussian and SE-adaLN-Zero. Beyond class-to-image generation, we also evaluate the generalization of the two improved methods on text-to-image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。