arXiv:2503.00245cs.LGcs.CL2025-03被引 3

让小模型也能用稀疏专家网络,提升手机端推理质量与效率。

CoSMoEs: Compact Sparse Mixture of Experts

  • 采用权重分解专家结构,提升小规模MoE模型性能。
  • 在同等计算量下,MoE模型比稠密模型准确率更高。
  • 优化模型卸载策略,显著降低手机端推理延迟。

稀疏专家混合(MoE)模型在大规模场景中广受欢迎,但在小型设备上仍研究不足。本文提出紧凑型稀疏专家混合(CoSMoEs),适用于设备端推理。针对设备端的三大挑战——性能、内存和延迟,我们证明在公平评估下(消除干扰因素),MoE架构在设备端的性能优于等计算量的稠密模型。通过引入权重分解专家,进一步提升模型表现。同时,显著提高模型卸载效率,从而降低推理延迟。

原文摘要 · Abstract (English)

Sparse Mixture of Expert (MoE) models are popular foundational architectures at large scale, however, under-explored at smaller sizes. Here, we show how to enable Compact Sparse Mixture of Experts (CoSMoEs) for on-device inference. Specifically, we tackle the three main on-device dimensions: Quality, Memory and Latency. Along the quality axis, we show that in a fair evaluation (removing confounding factors) MoE architectures outperform FLOP-aligned dense models at on-device scale. We introduce weight-decomposed experts, further improving the MoE model performance. Regarding model memory and latency, we significantly improve model offloading efficiency and, in turn, reduce model inference latency.

MoE小模型设备端推理稀疏专家

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。