首次为OPTQ和Qronos提供量化误差的理论保证,解释其设计合理性。
Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos
- 通过分析迭代过程,给出2-范数误差界,依赖校准数据与正则化参数。
- 证明随机版本可控制无穷范数,利于下游层和非线性层的量化。
- 首次为Qronos提供理论支持,解释其相比其他方法的实证优势。
后训练量化(PTQ)已成为降低现代深度神经网络(包括大语言模型)内存与计算成本的关键技术。在众多PTQ算法中,OPTQ框架(又称GPTQ)因其计算高效性和优异的实证表现成为领先方法。然而,该方法缺乏严格的定量理论保障。本文首次为OPTQ的确定性与随机变体,以及近期先进的相关算法Qronos,提供了定量误差边界。我们分析了OPTQ迭代过程引起的量化误差,推导出显式依赖于校准数据和正则化参数的非渐近2-范数误差界。该分析为多项实用设计选择(如按范数降序排列特征)提供了理论依据,并指导正则化参数的选择。对于随机变体,我们建立了更强的无穷范数误差界,可用于控制所需量化字母表大小,特别适用于下游层和非线性部分。最后,我们将分析扩展至Qronos,为其中确定性与随机变体提供了新的理论边界,有助于解释其经验优势。
原文摘要 · Abstract (English)
Post-training quantization (PTQ) has become a crucial tool for reducing the memory and compute costs of modern deep neural networks, including large language models (LLMs). Among PTQ algorithms, the OPTQ framework-also known as GPTQ-has emerged as a leading method due to its computational efficiency and strong empirical performance. Despite its widespread adoption, however, OPTQ lacks rigorous quantitative theoretical guarantees. This paper presents the first quantitative error bounds for both deterministic and stochastic variants of OPTQ, as well as for Qronos, a recent related state-of-the-art PTQ algorithm. We analyze how OPTQ's iterative procedure induces quantization error and derive non-asymptotic 2-norm error bounds that depend explicitly on the calibration data and a regularization parameter that OPTQ uses. Our analysis provides theoretical justification for several practical design choices, including the widely used heuristic of ordering features by decreasing norm, as well as guidance for selecting the regularization parameter. For the stochastic variant, we establish stronger infinity-norm error bounds, which enable control over the required quantization alphabet and are particularly useful for downstream layers and nonlinearities. Finally, we extend our analysis to Qronos, providing new theoretical bounds, for both its deterministic and stochastic variants, that help explain its empirical advantages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。