系统调研70个小型语言模型,揭示其性能与部署成本
Small Language Models: Survey, Measurements, and Insights
- 聚焦1亿至50亿参数的Transformer解码器模型,分析架构、数据集与训练算法创新
- 在常识推理、数学能力、上下文学习等任务中评估模型表现,覆盖长文本处理能力
- 实测推理延迟与内存占用,为设备端部署提供关键优化参考
小型语言模型(SLMs)虽广泛应用于现代智能设备,但相比主要部署于数据中心的大型语言模型(LLM),学术关注度显著不足。尽管研究者持续提升LLM能力以追求通用人工智能,SLM研究则致力于让机器智能更易用、经济和高效。本文聚焦于1亿至50亿参数的基于Transformer的解码器语言模型,系统调研70个最先进的开源SLM,从架构、训练数据集和训练算法三个维度分析其技术革新。同时,在常识推理、数学能力、上下文学习及长文本处理等多领域评估模型性能。为进一步理解其设备端运行开销,我们实测了推理延迟与内存占用。通过对基准测试数据的深入分析,提出推动该领域发展的关键洞见。
原文摘要 · Abstract (English)
Small language models (SLMs), despite their widespread adoption in modern smart devices, have received significantly less academic attention compared to their large language model (LLM) counterparts, which are predominantly deployed in data centers and cloud environments. While researchers continue to improve the capabilities of LLMs in the pursuit of artificial general intelligence, SLM research aims to make machine intelligence more accessible, affordable, and efficient for everyday tasks. Focusing on transformer-based, decoder-only language models with 100M-5B parameters, we survey 70 state-of-the-art open-source SLMs, analyzing their technical innovations across three axes: architectures, training datasets, and training algorithms. In addition, we evaluate their capabilities in various domains, including commonsense reasoning, mathematics, in-context learning, and long context. To gain further insight into their on-device runtime costs, we benchmark their inference latency and memory footprints. Through in-depth analysis of our benchmarking data, we offer valuable insights to advance research in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。