小模型如何在大模型时代突围:高效、私密、定制化
A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
- 以任务专精与资源受限为标准定义小语言模型,统一研究边界
- 提出小模型在隐私保护、低延迟场景下的应用优势与优化框架
- 适合关注轻量化部署、垂直领域适配与模型可信性的开发者
大语言模型(LLM)在文本生成、问答和推理等任务中展现出涌现能力,但如PaLM 540B和Llama-3.1 405B这类模型因参数量庞大,带来高计算开销,通常需依赖云API,引发隐私问题,限制边缘设备实时应用,并增加微调成本。此外,它们在医疗、法律等专业领域表现不佳,因缺乏领域知识,亟需专用模型。小语言模型(SLMs)因其低推理延迟、低成本、开发高效、易定制和适应性强而备受青睐,尤其适用于资源受限环境和领域知识获取,能有效解决隐私、响应速度和轻量微调需求。当前对SLMs的定义、获取、应用、增强及可靠性研究尚不系统,本综述针对这些方面构建分类体系与通用框架,提出以完成特定任务与适应资源受限场景作为定义标准,设定涌现能力最小规模与资源可承受最大规模的边界,推动小模型技术发展。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated emergent abilities in text generation, question answering, and reasoning, facilitating various tasks and domains. Despite their proficiency in various tasks, LLMs like PaLM 540B and Llama-3.1 405B face limitations due to large parameter sizes and computational demands, often requiring cloud API use which raises privacy concerns, limits real-time applications on edge devices, and increases fine-tuning costs. Additionally, LLMs often underperform in specialized domains such as healthcare and law due to insufficient domain-specific knowledge, necessitating specialized models. Therefore, Small Language Models (SLMs) are increasingly favored for their low inference latency, cost-effectiveness, efficient development, and easy customization and adaptability. These models are particularly well-suited for resource-limited environments and domain knowledge acquisition, addressing LLMs' challenges and proving ideal for applications that require localized data handling for privacy, minimal inference latency for efficiency, and domain knowledge acquisition through lightweight fine-tuning. The rising demand for SLMs has spurred extensive research and development. However, a comprehensive survey investigating issues related to the definition, acquisition, application, enhancement, and reliability of SLM remains lacking, prompting us to conduct a detailed survey on these topics. The definition of SLMs varies widely, thus to standardize, we propose defining SLMs by their capability to perform specialized tasks and suitability for resource-constrained settings, setting boundaries based on the minimal size for emergent abilities and the maximum size sustainable under resource constraints. For other aspects, we provide a taxonomy of relevant models/methods and develop general frameworks for each category to enhance and utilize SLMs effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。