开源16亿参数小模型Fox-1,兼顾性能与效率,适合云端边缘部署。
Fox-1: Open Small Language Model for Cloud and Edge
- 采用三阶段数据课程和8K序列长度,提升预训练效率。
- 在多个基准测试中超越或持平主流小模型,推理速度更快。
- 支持开源免费使用,适合研究者与开发者快速部署应用。
我们提出Fox-1系列小型语言模型,包括Fox-1-1.6B和Fox-1-1.6B-Instruct-v0.1。模型在3万亿个标记的网页文档数据上进行预训练,并用50亿个标记的指令跟随与多轮对话数据微调。为提升预训练效率,Fox-1-1.6B采用跨全部训练数据的三阶段数据课程,支持2K-8K序列长度。架构上,模型具备更深层数、更大词表和分组查询注意力(GQA),相比其他小模型更具性能与效率优势。在多个基准测试中,其表现优于或持平StableLM-2-1.6B、Gemma-2B、Qwen1.5-1.8B和OpenELM1.1B,同时具备优异的推理速度与吞吐量。模型权重以Apache 2.0许可证开源,旨在推动大模型民主化,向整个开源社区全面开放。
原文摘要 · Abstract (English)
We present Fox-1, a series of small language models (SLMs) consisting of Fox-1-1.6B and Fox-1-1.6B-Instruct-v0.1. These models are pre-trained on 3 trillion tokens of web-scraped document data and fine-tuned with 5 billion tokens of instruction-following and multi-turn conversation data. Aiming to improve the pre-training efficiency, Fox-1-1.6B model introduces a novel 3-stage data curriculum across all the training data with 2K-8K sequence length. In architecture design, Fox-1 features a deeper layer structure, an expanded vocabulary, and utilizes Grouped Query Attention (GQA), offering a performant and efficient architecture compared to other SLMs. Fox-1 achieves better or on-par performance in various benchmarks compared to StableLM-2-1.6B, Gemma-2B, Qwen1.5-1.8B, and OpenELM1.1B, with competitive inference speed and throughput. The model weights have been released under the Apache 2.0 license, where we aim to promote the democratization of LLMs and make them fully accessible to the whole open-source community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。