arXiv:2411.15242cs.LGcs.AI2024-11被引 21

Zamba2系列模型在保持顶尖性能的同时,大幅降低推理延迟与内存占用。

The Zamba2 Suite: Technical Report

  • 混合Mamba2-Transformer架构,优化训练数据与训练策略
  • 7.4B模型在三万亿令牌上训练,推理速度比同类模型快30%以上
  • 开源全部权重与预训练数据集,适合研究与部署场景

本文介绍Zamba2系列——一组1.2B、2.7B和7.4B参数的混合Mamba2-Transformer模型,在同规模开放权重模型中达到顶尖性能,同时显著提升推理延迟、吞吐量和内存效率。该系列基于Zamba1-7B的工作,通过优化架构、训练与退火数据集,并在高达三万亿令牌的数据上进行训练。我们公开了所有Zamba2模型的权重,以及指令微调版本,其表现可与同类指令模型媲美。此外,我们还开源了用于训练的预训练数据集Zyda-2。相关模型与数据集已开放获取:https://huggingface.co/Zyphra

原文摘要 · Abstract (English)

In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance against the leading open-weights models of their class, while achieving substantial gains in inference latency, throughput, and memory efficiency. The Zamba2 series builds upon our initial work with Zamba1-7B, optimizing its architecture, training and annealing datasets, and training for up to three trillion tokens. We provide open-source weights for all models of the Zamba2 series as well as instruction-tuned variants that are strongly competitive against comparable instruct-tuned models of their class. We additionally open-source the pretraining dataset, which we call Zyda-2, used to train the Zamba2 series of models. The models and datasets used in this work are openly available at https://huggingface.co/Zyphra

大模型混合架构开源高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。