PQuantML一键压缩模型,边端部署更高效
PQuantML: A Tool for End-to-End Hardware-aware Model Compression

- 统一接口支持剪枝与量化,可单独或联合使用
- 在喷注分类任务中实现参数量大幅减少,精度基本不变
- 专为边缘设备设计,适合实时高能物理数据处理
PQuantML 是一个开源的硬件感知神经网络模型压缩库,支持端到端工作流。针对严格延迟约束环境中的模型部署需求,该工具通过统一接口简化压缩模型的训练,支持多种粒度的剪枝方法以及固定精度量化,包含高粒度量化支持。在典型的喷注分类任务(即喷注标记)——一种与实时大型强子对撞机数据处理相关的边缘计算问题上进行了评估。结合不同剪枝方法与固定点量化,PQuantML 实现了显著的参数量和位宽降低,同时保持模型精度。压缩效果与 QKeras、HGQ 等现有工具进行对比,验证了其有效性。
原文摘要 · Abstract (English)
PQuantML is a new open-source, hardware-aware neural network model compression library tailored to end-to-end workflows. Motivated by the need to deploy performant models to environments with strict latency constraints, PQuantML simplifies training of compressed models by providing a unified interface to apply pruning and quantization, either jointly or individually. The library implements multiple pruning methods with different granularities, as well as fixed-point quantization with support for High-Granularity Quantization. We evaluate PQuantML on representative tasks such as the jet substructure classification, so-called jet tagging, an on-edge problem related to real-time LHC data processing. Using various pruning methods with fixed-point quantization, PQuantML achieves substantial parameter and bit-width reductions while maintaining accuracy. The resulting compression is further compared against existing tools, such as QKeras and HGQ.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。