通过蒸馏与量化扩展大模型家族,适配多样硬件需求。
Apertus LLM Family Expansion via Distillation and Quantization
- 基于8B模型蒸馏出4B参数的系列模型
- 在1.7万亿合规令牌上训练,保持强性能
- 适合资源受限场景下的高效部署
大模型的广泛应用催生了对多样化硬件和预算约束的适应需求。为此,我们验证了蒸馏与量化作为低成本扩展模型家族的有效手段。基于开源的Apertus 8B LLM,我们构建了Apertus-v1.1——一个最大达4B参数、在1.7万亿授权许可数据上训练的蒸馏模型家族。实验表明,该方法在覆盖广泛硬件与系统需求时兼具成本效益与优异的准确率表现。
原文摘要 · Abstract (English)
The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating the need for the models to satisfy certain budget and hardware constraints. This has led to the trend of LLMs being released in batches consisting of similar models of various sizes for the family of models to adhere to as wide of a range of constraints as possible. In this paper, we validate distillation and quantization as a cost-effective way to expand model families to new sizes and hardware formats. Based on the open-recipe Apertus 8B LLM, we produce Apertus-v1.1 - a distilled family of models with up to 4B parameters trained on 1.7T permissive license tokens. We demonstrate cost-efficiency and strong accuracy performance of our approach for covering large ranges of hardware and systems requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。