如何让大模型在手机等设备上高效运行?
On-Device Language Models: A Comprehensive Review
- 通过参数共享、模块化设计压缩模型体积
- 量化、剪枝、知识蒸馏使模型更轻量
- 适合移动开发者与边缘计算研究者参考
大型语言模型(LLMs)革新了自然语言处理应用,将LLMs部署在边缘设备上因降低延迟、实现数据本地化和个性化体验而愈发吸引人。本文全面回顾了在资源受限设备上部署计算密集型LLMs所面临的挑战,并探讨了多领域的创新解决方案。研究涵盖设备端语言模型的发展、高效架构(如参数共享与模块化设计),以及当前最先进的压缩技术,包括量化、剪枝和知识蒸馏。同时分析了硬件加速策略与协同边缘-云部署方式,揭示性能与资源利用之间的复杂权衡。来自主流手机厂商的设备端语言模型案例展示了实际应用与潜在优势。此外还讨论了自适应学习、多模态能力与个性化等关键议题。通过识别核心研究方向与开放挑战,本文为未来设备端语言模型发展提供了路线图,强调需跨学科协作以实现普适智能计算的潜力,同时保障负责任与伦理化的部署。更多研究工作与教育资源请访问 https://github.com/NexaAI/Awesome-LLMs-on-device。下载与运行设备端LLM请访问 https://www.nexaai.com/models。
原文摘要 · Abstract (English)
The advent of large language models (LLMs) revolutionized natural language processing applications, and running LLMs on edge devices has become increasingly attractive for reasons including reduced latency, data localization, and personalized user experiences. This comprehensive review examines the challenges of deploying computationally expensive LLMs on resource-constrained devices and explores innovative solutions across multiple domains. The paper investigates the development of on-device language models, their efficient architectures, including parameter sharing and modular designs, as well as state-of-the-art compression techniques like quantization, pruning, and knowledge distillation. Hardware acceleration strategies and collaborative edge-cloud deployment approaches are analyzed, highlighting the intricate balance between performance and resource utilization. Case studies of on-device language models from major mobile manufacturers demonstrate real-world applications and potential benefits. The review also addresses critical aspects such as adaptive learning, multi-modal capabilities, and personalization. By identifying key research directions and open challenges, this paper provides a roadmap for future advancements in on-device language models, emphasizing the need for interdisciplinary efforts to realize the full potential of ubiquitous, intelligent computing while ensuring responsible and ethical deployment. For a comprehensive review of research work and educational resources on on-device large language models (LLMs), please visit https://github.com/NexaAI/Awesome-LLMs-on-device. To download and run on-device LLMs, visit https://www.nexaai.com/models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。