扩散变压器自生表征引导,无需外部组件即可加速生成训练。
No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- 利用扩散变压器内部表征的渐进对齐实现自引导学习
- 在DiTs和SiTs上提升性能,超越依赖辅助任务的方法
- 适合追求高效生成模型训练的研究者与开发者
近期研究表明,学习有意义的内部表征可加速生成模型训练。然而,现有方法要么引入现成的外部表征任务,要么依赖大规模预训练的外部编码器提供表征引导。本文提出自表征对齐(SRA),一种简单有效的方法,仅通过已训练扩散变压器的内部表征获取引导信号。SRA将高噪声条件下早期层的隐含表示与低噪声条件下后期层的表示对齐,从而在训练过程中逐步增强整体表征学习。实验表明,SRA应用于DiTs和SiTs均带来一致性能提升,显著优于依赖辅助表征任务的方法,其性能接近依赖外部预训练编码器的方法,验证了扩散变压器自身具备表征对齐加速训练的可行性。
原文摘要 · Abstract (English)
Recent studies have demonstrated that learning a meaningful internal representation can accelerate generative training. However, existing approaches necessitate to either introduce an off-the-shelf external representation task or rely on a large-scale, pre-trained external representation encoder to provide representation guidance during the training process. In this study, we posit that the unique discriminative process inherent to diffusion transformers enables them to offer such guidance without requiring external representation components. We propose SelfRepresentation Alignment (SRA), a simple yet effective method that obtains representation guidance using the internal representations of learned diffusion transformer. SRA aligns the latent representation of the diffusion transformer in the earlier layer conditioned on higher noise to that in the later layer conditioned on lower noise to progressively enhance the overall representation learning during only the training process. Experimental results indicate that applying SRA to DiTs and SiTs yields consistent performance improvements, and largely outperforms approaches relying on auxiliary representation task. Our approach achieves performance comparable to methods that are dependent on an external pre-trained representation encoder, which demonstrates the feasibility of acceleration with representation alignment in diffusion transformers themselves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。