统一量化与剪枝的随机框架,实现1比特压缩与理论保障
Unified Stochastic Framework for Neural Network Quantization and Pruning
- 用随机路径追踪统一处理量化与剪枝
- 支持1比特量化,误差有严格理论保证
- 适合需要高压缩比与可靠性的模型部署场景
量化与剪枝是压缩神经网络的两种关键技术,但通常被独立处理,缺乏理论关联。本文提出一种基于随机路径追踪算法的统一后训练压缩框架,扩展了SPFQ方法的应用范围至剪枝和低比特量化(包括1比特)。通过引入缩放参数并泛化随机算子,该方法具备鲁棒误差校正能力,并为量化、剪枝及其组合提供了严格的理论误差界。
原文摘要 · Abstract (English)
Quantization and pruning are two essential techniques for compressing neural networks, yet they are often treated independently, with limited theoretical analysis connecting them. This paper introduces a unified framework for post-training quantization and pruning using stochastic path-following algorithms. Our approach builds on the Stochastic Path Following Quantization (SPFQ) method, extending its applicability to pruning and low-bit quantization, including challenging 1-bit regimes. By incorporating a scaling parameter and generalizing the stochastic operator, the proposed method achieves robust error correction and yields rigorous theoretical error bounds for both quantization and pruning as well as their combination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。