arXiv:2410.02098cs.CVcs.LG2024-10ICLR被引 32

让扩散模型按需分配算力,实现970亿参数高效生成

EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing

  • 采用专家选择路由机制,动态分配计算资源
  • 970亿参数模型训练收敛更快,生成质量显著提升
  • 适合追求高精度与推理效率的图像生成研究者

扩散变压器在文本到图像生成中广泛应用。尽管将模型扩展至数十亿参数展现出潜力,但超越当前规模的有效性仍缺乏探索且面临挑战。通过显式利用图像生成中的计算异构性,我们提出一种新型混合专家(MoE)模型EC-DIT,用于扩散变压器,其采用专家选择路由机制。该方法可自适应优化理解输入文本和生成图像块的计算分配,使计算异构性与文本-图像复杂度变化相匹配。这种异构性使EC-DIT得以扩展至970亿参数,并在训练收敛速度、文本-图像对齐性和整体生成质量上优于密集模型和传统MoE模型。通过大量消融实验表明,EC-DIT能通过端到端训练识别不同文本重要性,展现卓越可扩展性与自适应算力分配能力。值得注意的是,在文本到图像对齐评估中,最大模型达到71.68%的GenEval得分,同时保持具有竞争力的推理速度与直观可解释性。

原文摘要 · Abstract (English)

Diffusion transformers have been widely adopted for text-to-image synthesis. While scaling these models up to billions of parameters shows promise, the effectiveness of scaling beyond current sizes remains underexplored and challenging. By explicitly exploiting the computational heterogeneity of image generations, we develop a new family of Mixture-of-Experts (MoE) models (EC-DIT) for diffusion transformers with expert-choice routing. EC-DIT learns to adaptively optimize the compute allocated to understand the input texts and generate the respective image patches, enabling heterogeneous computation aligned with varying text-image complexities. This heterogeneity provides an efficient way of scaling EC-DIT up to 97 billion parameters and achieving significant improvements in training convergence, text-to-image alignment, and overall generation quality over dense models and conventional MoE models. Through extensive ablations, we show that EC-DIT demonstrates superior scalability and adaptive compute allocation by recognizing varying textual importance through end-to-end training. Notably, in text-to-image alignment evaluation, our largest models achieve a state-of-the-art GenEval score of 71.68% and still maintain competitive inference speed with intuitive interpretability.

扩散模型MoE文本生成高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。