为专家混合模型提出新参数化,实现超大规模模型的高效调参
$μ$-Parametrization for Mixture of Experts
- 基于μ-参数化理论,建立专家混合模型的宽度泛化机制
- 实验验证最优学习率可跨模型规模可靠迁移,降低调参成本
- 适合研究超大模型训练与高效调参的团队参考
近年来,大型语言模型(LLM)发展迅速,专家混合(MoE)架构成为超大规模模型的主流设计。当前最大的开源模型已超过1万亿参数。在如此规模下,超参数调优成本极高。为此,μTransfer技术逐渐成为关键方法,可在不同模型尺度间无缝迁移最优超参数,显著降低调优开销。然而,现有研究主要聚焦于密集型模型,对MoE架构探索不足。本文推导出适用于MoE的μ-参数化,为模型宽度变化下的特征学习提供理论保障。实验表明,最优学习率可在不同模型规模间稳定迁移,为大规模MoE模型的高效超参数调优奠定基础。
原文摘要 · Abstract (English)
Recent years have seen a growing interest and adoption of LLMs, with Mixture-of-Experts (MoE) emerging as a leading architecture in extremely large models. Currently, the largest open-source models reach over $1$T parameters. At such scales, hyperparameter tuning becomes prohibitively expensive. Precisely for this reason, the $μ$Transfer is becoming a key technique. It allows for seamless transfer of optimal hyperparameters across model scales, resulting in a huge reduction in tuning costs. However, existing work has primarily focused on dense LLMs, leaving MoE architectures unexplored. In this work, we derive a $μ$-Parameterization for MoE, providing theoretical guarantees for feature learning across model widths. Our experiments demonstrate that the optimal learning rate reliably transfers across model sizes, establishing a foundation for efficient hyperparameter tuning in large-scale MoE models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。