提出MoE-CAP基准,揭示稀疏专家模型在成本、精度与性能间的权衡难题
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
- 构建针对稀疏专家系统的专用评估框架,揭示三者不可兼得的权衡规律
- 发现当前硬件下难以同时优化成本、精度与性能,普遍存在二选一困境
- 引入新指标S-MBU与S-MFU,适配多平台部署的稀疏模型性能评估
稀疏混合专家(MoE)架构正被广泛用于高效扩展大语言模型,但其依赖异构计算与内存资源,共同影响系统成本、精度与性能(CAP),导致不可避免的权衡。现有基准常无法准确捕捉这些权衡,阻碍实际部署决策。为此,我们提出MoE-CAP,一个专为MoE系统设计的基准。分析表明,当前硬件下难以实现CAP三者的最优平衡,系统通常在其中两项上优化而牺牲第三项——这一现象被称为MoE-CAP权衡。为可视化该关系,我们提出CAP雷达图。此外,我们引入稀疏感知性能指标:稀疏内存带宽利用率(S-MBU)与稀疏模型浮点运算利用率(S-MFU),以实现跨多种硬件平台与部署场景的精准性能评估。
原文摘要 · Abstract (English)
The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory resources. These factors jointly affect system Cost, Accuracy, and Performance (CAP), making trade-offs inevitable. Existing benchmarks often fail to capture these trade-offs accurately, complicating practical deployment decisions. To address this, we introduce MoE-CAP, a benchmark specifically designed for MoE systems. Our analysis reveals that achieving an optimal balance across CAP is difficult with current hardware; MoE systems typically optimize two of the three dimensions at the expense of the third-a dynamic we term the MoE-CAP trade-off. To visualize this, we propose the CAP Radar Diagram. We further introduce sparsity-aware performance metrics-Sparse Memory Bandwidth Utilization (S-MBU) and Sparse Model FLOPS Utilization (S-MFU)-to enable accurate performance benchmarking of MoE systems across diverse hardware platforms and deployment scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。