首个系统性评测后训练稀疏化的基准,助力高效模型设计
PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models
- 构建覆盖40+模型的可插拔稀疏化评测框架
- 验证10余种算法在3类任务上的性能与稀疏能力
- 为算法优化和模型设计提供实证指导
随着对模型效率关注增加,后训练稀疏化(PTS)因其高效性日益普及。然而,当前仍缺乏对PTS算法最佳实践及模型稀疏化能力的系统评估,制约该领域发展。为此,本文提出首个面向算法与模型的综合性后训练稀疏化基准PTS-Bench。在超过40个现成模型架构上,针对3类典型任务,对10余种通用、细粒度的可插拔稀疏化技术进行系统评测。通过大量实验与分析,获得多项有价值结论,从算法与模型双视角提供深入洞察。本基准可为理解PTS算法提供新发现,全面评估模型稀疏化能力,并提供结构清晰、易集成的开源框架。我们期望此工作能为未来后训练稀疏化方法研究与稀疏友好模型设计提供重要参考。代码已公开于https://github.com/ModelTC/msbench。
原文摘要 · Abstract (English)
With the increased attention to model efficiency, post-training sparsity (PTS) has become more and more prevalent because of its effectiveness and efficiency. However, there remain questions on better practice of PTS algorithms and the sparsification ability of models, which hinders the further development of this area. Therefore, a benchmark to comprehensively investigate the issues above is urgently needed. In this paper, we propose the first comprehensive post-training sparsity benchmark called PTSBench towards algorithms and models. We benchmark 10+ PTS general-pluggable fine-grained techniques on 3 typical tasks using over 40 off-the-shelf model architectures. Through extensive experiments and analyses, we obtain valuable conclusions and provide several insights from both algorithms and model aspects. Our PTSBench can provide (1) new observations for a better understanding of the PTS algorithms, (2) in-depth and comprehensive evaluations for the sparsification ability of models, and (3) a well-structured and easy-integrate open-source framework. We hope this work will provide illuminating conclusions and advice for future studies of post-training sparsity methods and sparsification-friendly model design. The code for our PTSBench is released at \href{https://github.com/ModelTC/msbench}{https://github.com/ModelTC/msbench}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。