首次系统分析低比特量化对高维线性回归的影响,揭示其噪声与误差机制。
Learning under Quantization for High-Dimensional Linear Regression
- 构建理论框架,分析数据、参数、梯度等五类量化对学习的影响
- 发现乘性量化比加性量化更少破坏数据谱结构,降低误差
- 为硬件受限下的模型训练提供理论指导,适合算法与系统研究者
低比特量化已成为大规模模型高效训练的关键技术。尽管在实践中广泛应用,其对学习性能的理论理解仍不充分,甚至在最简单的线性回归场景中也缺乏系统研究。本文首次对高维线性回归下有限步随机梯度下降(SGD)中的多种量化目标(数据、标签、参数、激活、梯度)进行系统性理论分析。提出新颖的解析框架,建立算法依赖和数据依赖的过量风险边界,揭示不同量化方式的影响:参数、激活和梯度量化会放大训练噪声;数据量化则扭曲数据谱并引入额外近似误差。关键发现:加性量化(恒定步长)的噪声放大受批量大小抑制,而乘性量化(输入相关步长)能较好保留谱结构,减少谱失真。在常见多项式衰减数据谱下,定量比较了乘性与加性量化风险,类比于浮点与整数量化方法的差异。该理论为理解量化如何塑造优化算法的学习动态提供了有力视角,推动在实际硬件约束下的学习理论发展。
原文摘要 · Abstract (English)
The use of low-bit quantization has emerged as an indispensable technique for enabling the efficient training of large-scale models. Despite its widespread empirical success, a rigorous theoretical understanding of its impact on learning performance remains notably absent, even in the simplest linear regression setting. We present the first systematic theoretical study of this fundamental question, analyzing finite-step stochastic gradient descent (SGD) for high-dimensional linear regression under a comprehensive range of quantization targets: data, label, parameter, activation, and gradient. Our novel analytical framework establishes precise algorithm-dependent and data-dependent excess risk bounds that characterize how different quantization affects learning: parameter, activation, and gradient quantization amplify noise during training; data quantization distorts the data spectrum and introduces additional approximation error. Crucially, we distinguish the effects of two quantization schemes: we prove that for additive quantization (with constant quantization steps), the noise amplification benefits from a suppression effect scaled by the batch size, while multiplicative quantization (with input-dependent quantization steps) largely preserves the spectral structure, thereby reducing the spectral distortion. Furthermore, under common polynomial-decay data spectra, we quantitatively compare the risks of multiplicative and additive quantization, drawing a parallel to the comparison between FP and integer quantization methods. Our theory provides a powerful lens to characterize how quantization shapes the learning dynamics of optimization algorithms, paving the way to further explore learning theory under practical hardware constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。