综述小型语言模型的架构与优化技术,助力高效部署。
A Survey of Small Language Models
- 提出新分类体系,涵盖压缩、剪枝、量化等优化方法
- 梳理常用评估数据集与指标,支持模型对比研究
- 面向研发人员,提供小模型落地的技术指南
小型语言模型(SLMs)因高效性与低资源消耗,日益成为执行各类语言任务的理想选择,尤其适用于设备端、移动端及边缘计算场景。本文全面综述了SLMs在模型架构、训练方法及模型压缩技术方面的进展。提出一种新型分类体系,用于归纳优化SLMs的方法,包括模型压缩、剪枝和量化技术。总结了用于评估SLMs的基准数据集及常用评价指标。同时,指出了当前仍需解决的关键开放挑战。本综述旨在为关注小型高效语言模型研发与部署的研究者与实践者提供有价值的参考。
原文摘要 · Abstract (English)
Small Language Models (SLMs) have become increasingly important due to their efficiency and performance to perform various language tasks with minimal computational resources, making them ideal for various settings including on-device, mobile, edge devices, among many others. In this article, we present a comprehensive survey on SLMs, focusing on their architectures, training techniques, and model compression techniques. We propose a novel taxonomy for categorizing the methods used to optimize SLMs, including model compression, pruning, and quantization techniques. We summarize the benchmark datasets that are useful for benchmarking SLMs along with the evaluation metrics commonly used. Additionally, we highlight key open challenges that remain to be addressed. Our survey aims to serve as a valuable resource for researchers and practitioners interested in developing and deploying small yet efficient language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。