arXiv:2507.08836cs.LGcs.PF2025-07被引 1

CompactifAI压缩Llama 3.1 8B模型,兼顾效率与精度。

Accuracy and Consumption analysis from a compressed model by CompactifAI from Multiverse Computing

  • 用CompactifAI对Llama 3.1 8B进行压缩,降低计算开销。
  • 压缩后模型能耗显著下降,精度保持稳定。
  • 适合关注模型轻量化部署的开发者与研究者。

本研究评估了Multiverse Computing开发的压缩方法CompactifAI在大型语言模型Llama 3.1 8B上的表现。通过Codecarbon框架衡量能耗,Ragas框架评估准确性,对比了使用CompactifAI压缩后的模型与原始全尺寸版本。结果表明,压缩模型不仅显著降低了计算资源消耗,同时维持了原有精度,提升了模型的效率、可扩展性与成本效益。

原文摘要 · Abstract (English)

This study evaluates the performance of a compression method, called CompactifAI, developed by Multiverse Computing, applied to the large language model Llama 3.1 8B\cite{llama}. The evaluation focused on model efficiency (in terms of energy consumption) and accuracy using respectively the frameworks Codecarbon\cite{codecarbon} and Ragas\cite{ragas}. A comparison was performed between the model compressed with CompactifAI\cite{compactifai}\cite{compactifai2} and its full-size version. Our findings reveal that the compressed model using CompactifAI not only significantly reduced the computational resources but also maintained the model accuracy, making the model more efficient, scalable and cost-effective.

模型压缩Llama 3能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。