arXiv:2410.21352cs.CLcs.AI2024-10NeurIPS被引 24

构建首个通用大模型压缩评估基准,指导实际部署选择最佳压缩方案。

LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment

  • 设计多维度评估体系,覆盖主流压缩方法与真实场景需求
  • 在多个模型、数据集上验证,提供可复现的性能对比结果
  • 适合研究人员和工程师选型压缩技术,推动高效部署落地

尽管大语言模型展现出强大智能能力,但其对计算与存储的高需求限制了实际应用。为此,众多模型压缩技术被提出以提升效率。然而,现有研究仅在有限模型、数据集和指标上验证方法,缺乏在更广泛场景下的全面评估。因此,在特定情况下应采用哪种压缩方法仍不明确。为弥合这一差距,我们提出大语言模型压缩基准(LLMCBench),一个精心设计的基准测试平台,包含深入分析的压缩算法评估体系。我们首先分析实际生产需求,设计评估路径与指标;随后使用多种主流压缩方法开展大规模实验与对比;最后基于评估结果进行深入分析,为压缩算法设计提供实用洞见。我们希望 LLMCBench 能为压缩算法设计提供有益建议,并成为未来研究的基础。代码已开源:https://github.com/AboveParadise/LLMCBench。

原文摘要 · Abstract (English)

Although large language models (LLMs) have demonstrated their strong intelligence ability, the high demand for computation and storage hinders their practical application. To this end, many model compression techniques are proposed to increase the efficiency of LLMs. However, current researches only validate their methods on limited models, datasets, metrics, etc, and still lack a comprehensive evaluation under more general scenarios. So it is still a question of which model compression approach we should use under a specific case. To mitigate this gap, we present the Large Language Model Compression Benchmark (LLMCBench), a rigorously designed benchmark with an in-depth analysis for LLM compression algorithms. We first analyze the actual model production requirements and carefully design evaluation tracks and metrics. Then, we conduct extensive experiments and comparison using multiple mainstream LLM compression approaches. Finally, we perform an in-depth analysis based on the evaluation and provide useful insight for LLM compression design. We hope our LLMCBench can contribute insightful suggestions for LLM compression algorithm design and serve as a foundation for future research. Our code is available at https://github.com/AboveParadise/LLMCBench.

大模型压缩评估基准模型部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。