系统梳理高效大模型架构,助力快速部署与资源节省
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
- 按线性、稀疏建模等思路优化传统Transformer结构
- 涵盖高效注意力、专家混合及混合架构等多种技术路径
- 适合关注模型轻量化与实际应用的研究者与开发者
大语言模型(LLMs)在语言理解、生成和多模态能力上表现卓越。以Transformer为基础的现代LLM虽具良好扩展性,但计算开销巨大,制约了大规模训练与实际部署。本文系统梳理了突破Transformer局限的创新架构:从语言建模出发,涵盖线性与稀疏序列建模、高效全注意力变体、稀疏专家混合(MoE)、融合多种技术的混合架构,以及新兴的扩散式LLM。同时讨论这些技术在多模态中的应用及其对可扩展、资源感知基础模型发展的意义。通过分类归纳近年研究,本综述为现代高效LLM架构提供蓝图,旨在推动更高效、通用的人工智能系统发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have delivered impressive results in language understanding, generation, reasoning, and pushes the ability boundary of multimodal models. Transformer models, as the foundation of modern LLMs, offer a strong baseline with excellent scaling properties. However, the traditional transformer architecture requires substantial computations and poses significant obstacles for large-scale training and practical deployment. In this survey, we offer a systematic examination of innovative LLM architectures that address the inherent limitations of transformers and boost the efficiency. Starting from language modeling, this survey covers the background and technical details of linear and sparse sequence modeling methods, efficient full attention variants, sparse mixture-of-experts, hybrid model architectures incorporating the above techniques, and emerging diffusion LLMs. Additionally, we discuss applications of these techniques to other modalities and consider their wider implications for developing scalable, resource-aware foundation models. By grouping recent studies into the above category, this survey presents a blueprint of modern efficient LLM architectures, and we hope this could help motivate future research toward more efficient, versatile AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。