提出新指标衡量扩散Transformer内部表征多样性,提升生成质量。
DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers

- 引入加权多样性分数(WDS)量化不同层间表征差异。
- WDS与生成质量强相关(皮尔逊r=-0.869),可作性能指标。
- 设计DiverseDiT++框架,通过残差连接和损失函数增强多样性。
扩散Transformer(DiTs)在视觉合成方面取得显著进展,得益于其优越的可扩展性。为提升DiTs捕捉有意义内部表征的能力,近期工作如REPA引入外部预训练编码器进行表征对齐。然而,社区对DiTs内部表征学习机制仍理解不足。本文首次系统分析DiTs的表征动态,通过量化块级表征多样性。我们提出一种新指标——加权多样性分数(WDS),用于测量不同层间表征差异。在多种设置下对内部表征演化与影响的广泛研究揭示:块间表征多样性是有效表征学习的关键因素。更重要的是,WDS与不同设置、模型规模和训练阶段下的生成质量高度相关(与log(FID)的皮尔逊相关系数r = -0.869),表明其具备作为模型性能指示器和优化指导的潜力。基于此发现,我们提出DiverseDiT++,一个显式促进表征多样性的新框架。具体而言,方法通过长残差连接多样化各层输入表征,并引入表征多样性损失以鼓励各层学习不同特征。在ImageNet 256×256和512×512上的大量实验表明,DiverseDiT++在不同规模的骨干网络上均带来一致的性能提升与收敛加速。
原文摘要 · Abstract (English)
Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs remain poorly understood in the community. To this end, this paper first presents a systematic analysis of the representation dynamics of DiTs via quantifying the diversity of block-wise representations. Specifically, we introduce a novel metric, termed the Weighted Diversity Score (WDS), to measure the representational discrepancies across different blocks. Through extensive investigations on the evolution and influence of internal representations under various settings, we reveal that representation diversity across blocks is a critical factor for effective representation learning in DiTs. More importantly, WDS exhibits a strong correlation with synthesis quality across diverse settings, model scales, and training stages (Pearson's $r=-0.869$ with $\log(\text{FID})$), suggesting its potential as an indicator to reflect model performance and a principled guide for model optimization. Based on this key finding, we propose DiverseDiT++, a novel framework that explicitly promotes diverse representation learning. Concretely, our method incorporates long residual connections to diversify input representations across blocks and a representation diversity loss to encourage blocks to learn distinct features. Extensive experiments on ImageNet $256\times256$ and $512\times512$ demonstrate that our DiverseDiT++ yields consistent performance gains and convergence acceleration when applied to different backbones with various sizes,...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。