arXiv:2603.04239cs.CV2026-03中稿 · CVPR被引 4

提升扩散Transformer的表征多样性,加速收敛并增强性能

DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers

  • 引入长残差连接与多样性损失,显式促进各层表征差异
  • 在ImageNet上实现一致性能提升,一阶段生成也有效
  • 适用于不同规模模型,可与现有方法互补

最近的扩散Transformer(DiTs)在视觉合成领域取得突破,因其出色的可扩展性。为增强DiTs捕捉有意义内部表征的能力,已有工作如REPA引入外部预训练编码器进行表征对齐,但其内部表征学习机制仍不清晰。为此,我们系统研究了DiTs的表征动态,发现块间表征多样性是有效学习的关键因素。基于此,提出DiverseDiT框架:通过长残差连接多样化各层输入表征,并引入表征多样性损失,促使各块学习不同特征。在ImageNet 256x256和512x512上的大量实验表明,该方法在不同规模骨干网络上均带来稳定性能提升与收敛加速,即使在挑战性的单步生成设置下也有效。此外,DiverseDiT与现有表征学习技术具有互补性,可进一步提升性能。本工作揭示了DiTs表征学习的动态规律,提供了一种实用高效的性能增强方案。

原文摘要 · Abstract (English)

Recent breakthroughs in Diffusion Transformers (DiTs) have revolutionized the field of visual synthesis due to their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs are not well understood. To this end, we first systematically investigate the representation dynamics of DiTs. Through analyzing the evolution and influence of internal representations under various settings, we reveal that representation diversity across blocks is a crucial factor for effective learning. Based on this key insight, we propose DiverseDiT, a novel framework that explicitly promotes representation diversity. DiverseDiT incorporates long residual connections to diversify input representations across blocks and a representation diversity loss to encourage blocks to learn distinct features. Extensive experiments on ImageNet 256x256 and 512x512 demonstrate that our DiverseDiT yields consistent performance gains and convergence acceleration when applied to different backbones with various sizes, even when tested on the challenging one-step generation setting. Furthermore, we show that DiverseDiT is complementary to existing representation learning techniques, leading to further performance gains. Our work provides valuable insights into the representation learning dynamics of DiTs and offers a practical approach for enhancing their performance.

扩散模型表征学习Transformer图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。