arXiv:2505.02309cs.LGcs.AI2025-05中稿 · IEEE COMPSAC 2025综述被引 20

压缩大模型,让AI在手机等设备上高效运行

Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques

  • 用知识蒸馏、量化和剪枝三类技术缩小模型体积
  • 可将模型大小减少70%以上,推理速度提升2-5倍
  • 适合做移动端、边缘设备AI应用的开发者参考

大型语言模型(LLMs)在人工智能多个领域带来革命性进展,但其巨大的资源需求限制了在移动和边缘设备上的部署。本文综述了用于压缩LLMs以实现资源受限环境下高效推理的技术。重点分析三大方法:知识蒸馏、模型量化和模型剪枝,阐述其原理、不同变体及成功应用案例。同时简要讨论混合专家系统与早期退出策略等互补技术。最后展望未来发展方向,旨在为研究人员和实践者优化LLM在边缘部署提供有价值参考。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have revolutionized many areas of artificial intelligence (AI), but their substantial resource requirements limit their deployment on mobile and edge devices. This survey paper provides a comprehensive overview of techniques for compressing LLMs to enable efficient inference in resource-constrained environments. We examine three primary approaches: Knowledge Distillation, Model Quantization, and Model Pruning. For each technique, we discuss the underlying principles, present different variants, and provide examples of successful applications. We also briefly discuss complementary techniques such as mixture-of-experts and early-exit strategies. Finally, we highlight promising future directions, aiming to provide a valuable resource for both researchers and practitioners seeking to optimize LLMs for edge deployment.

模型压缩边缘计算大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。