arXiv:2505.03531cs.CLcs.LG2025-05ACL被引 7

优化细粒度MoE模型推理,提升效率且不损失性能

Faster MoE LLM Inference for Extremely Large Models

  • 通过减少激活专家数量,显著提升推理效率
  • 在不降性能前提下,吞吐量提升至少10%
  • 适合大规模MoE模型部署与高效服务场景

稀疏混合专家(MoE)大语言模型正成为超大规模模型的主流架构。现有优化主要聚焦于粗粒度MoE结构。随着DeepSeek模型的出现,细粒度MoE逐渐流行,但相关研究仍有限。本文探讨不同服务负载下的推理效率动态,并研究减少激活专家数和总专家数对效率与性能权衡的影响。结果表明:减少激活专家数可在特定场景下带来显著效率提升,性能损失微小;而减少总专家数虽有少量效率改善,但导致严重性能下降。所提方法可在不损失性能前提下,将吞吐量提升至少10%。总体而言,MoE推理优化仍具巨大探索空间。

原文摘要 · Abstract (English)

Sparse Mixture of Experts (MoE) large language models (LLMs) are gradually becoming the mainstream approach for ultra-large-scale models. Existing optimization efforts for MoE models have focused primarily on coarse-grained MoE architectures. With the emergence of DeepSeek Models, fine-grained MoE models are gaining popularity, yet research on them remains limited. Therefore, we want to discuss the efficiency dynamic under different service loads. Additionally, fine-grained models allow deployers to reduce the number of routed experts, both activated counts and total counts, raising the question of how this reduction affects the trade-off between MoE efficiency and performance. Our findings indicate that while deploying MoE models presents greater challenges, it also offers significant optimization opportunities. Reducing the number of activated experts can lead to substantial efficiency improvements in certain scenarios, with only minor performance degradation. Reducing the total number of experts provides limited efficiency gains but results in severe performance degradation. Our method can increase throughput by at least 10\% without any performance degradation. Overall, we conclude that MoE inference optimization remains an area with substantial potential for exploration and improvement.

MoE模型推理优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。