arXiv:2501.05465cs.CL2025-01综述被引 1

小模型也能强:10亿参数内模型可媲美大模型性能

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)

  • 聚焦10亿到80亿参数的小型语言模型,探索其通用与专用场景表现
  • 发现部分小模型在多项任务上超越大型模型,证明规模非唯一路径
  • 适合追求高效、低成本部署的开发者与研究者参考

随着基础大模型持续增大,一个关键问题浮现:是否唯有大规模才是前进方向?本综述整合约160篇论文,系统梳理了参数量在10亿至80亿之间的小型语言模型(SLMs),揭示这些小模型在多项任务中表现不逊于甚至优于大型模型。研究涵盖无任务依赖的通用型SLMs、特定任务优化的SLMs,以及构建SLMs的关键技术,旨在为社区提供兼顾性能、效率、可扩展性与成本的建模指引。此外,本文定义并刻画了SLMs的‘有效规模’,量化其相对于大模型的实际能力提升。

原文摘要 · Abstract (English)

As foundation AI models continue to increase in size, an important question arises - is massive scale the only path forward? This survey of about 160 papers presents a family of Small Language Models (SLMs) in the 1 to 8 billion parameter range that demonstrate smaller models can perform as well, or even outperform large models. We explore task agnostic, general purpose SLMs, task-specific SLMs and techniques to create SLMs that can guide the community to build models while balancing performance, efficiency, scalability and cost. Furthermore we define and characterize SLMs' effective sizes, representing increased capability with respect to LLMs.

小模型语言模型效率优化综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。