量化模型训练中,权重量化范围影响学习动态与泛化误差变化模式。
High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
- 基于高维极限理论,推导出直通估计器的确定性微分方程动态
- 发现泛化误差先平台后骤降,平台期长度由量化范围决定
- 揭示量化模型与全精度模型的渐近偏差,适合研究量化训练机制者
量化神经网络训练优化的是离散且不可导的目标函数。直通估计器(STE)通过代理梯度实现反向传播,被广泛采用。以往研究多关注代理梯度性质与收敛性,但量化超参数(如位宽、量化范围)对学习动态的影响仍不明确。本文在高维极限下证明,STE动态收敛至确定性常微分方程。结果表明,STE训练呈现泛化误差先平台后骤降的行为,平台持续时间取决于量化范围。固定点分析量化了其与未量化线性模型的渐近偏差。此外,还将随机梯度下降的解析方法拓展至权重与输入的非线性变换。
原文摘要 · Abstract (English)
Quantized neural network training optimizes a discrete, non-differentiable objective. The straight-through estimator (STE) enables backpropagation through surrogate gradients and is widely used. While previous studies have primarily focused on the properties of surrogate gradients and their convergence, the influence of quantization hyperparameters, such as bit width and quantization range, on learning dynamics remains largely unexplored. We theoretically show that in the high-dimensional limit, STE dynamics converge to a deterministic ordinary differential equation. This reveals that STE training exhibits a plateau followed by a sharp drop in generalization error, with plateau length depending on the quantization range. A fixed-point analysis quantifies the asymptotic deviation from the unquantized linear model. We also extend analytical techniques for stochastic gradient descent to nonlinear transformations of weights and inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。