arXiv:2503.16057cs.CVcs.AI2025-03ICML被引 12

提出动态专家路由机制,提升扩散模型扩展性与生成质量。

Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts

  • 让令牌与专家竞争选优,动态分配关键任务给最优专家。
  • 在ImageNet上实现显著性能提升,且具备良好可扩展性。
  • 适合研究大规模视觉生成与专家混合模型的开发者。

扩散模型已成为视觉生成的主流框架。在此基础上,混合专家(MoE)方法展现出提升模型可扩展性和性能的潜力。本文提出Race-DiT,一种基于灵活路由策略Expert Race的扩散变换器MoE模型。通过让令牌与专家共同竞争并选择最优候选者,模型能够动态将专家分配给关键令牌。此外,我们引入逐层正则化以解决浅层学习难题,并设计路由器相似性损失防止模式崩溃,从而提升专家利用率。在ImageNet上的大量实验验证了该方法的有效性,不仅带来显著性能提升,还具备良好的可扩展性。

原文摘要 · Abstract (English)

Diffusion models have emerged as mainstream framework in visual generation. Building upon this success, the integration of Mixture of Experts (MoE) methods has shown promise in enhancing model scalability and performance. In this paper, we introduce Race-DiT, a novel MoE model for diffusion transformers with a flexible routing strategy, Expert Race. By allowing tokens and experts to compete together and select the top candidates, the model learns to dynamically assign experts to critical tokens. Additionally, we propose per-layer regularization to address challenges in shallow layer learning, and router similarity loss to prevent mode collapse, ensuring better expert utilization. Extensive experiments on ImageNet validate the effectiveness of our approach, showcasing significant performance gains while promising scaling properties.

扩散模型专家混合视觉生成路由机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。