arXiv:2602.19241stat.MLcs.AI2026-02被引 3

揭示低精度训练中模型与数据规模的理论关系,指导高效训练设计。

Scaling Laws for Precision in High-Dimensional Linear Regression

  • 在高维线性回归框架下分析量化对模型和数据容量的影响。
  • 乘法量化不降低有效模型规模,加法量化则会减少有效模型规模。
  • 为硬件受限场景下的训练优化提供理论依据,适合系统与算法研究者。

低精度训练对平衡模型质量与训练成本至关重要,需协同分配模型规模、数据集规模与数值精度。尽管经验性缩放定律表明量化会影响有效模型和数据容量,或表现为加性误差,但其背后的理论机制仍不清楚。本文首次在高维投影线性回归框架下开展低精度训练的理论研究。通过分析信号相关的乘法量化与信号无关的加法量化,我们发现二者存在关键差异:虽然两者均引入加性误差并降低有效数据规模,但乘法量化保持全精度模型规模,而加法量化会减小有效模型规模。数值实验验证了理论结果。本工作严谨刻画了模型规模、数据规模与量化误差之间的复杂相互作用,为在实际硬件约束下优化训练协议提供了理论基础。

原文摘要 · Abstract (English)

Low-precision training is critical for optimizing the trade-off between model quality and training costs, necessitating the joint allocation of model size, dataset size, and numerical precision. While empirical scaling laws suggest that quantization impacts effective model and data capacities or acts as an additive error, the theoretical mechanisms governing these effects remain largely unexplored. In this work, we initiate a theoretical study of scaling laws for low-precision training within a high-dimensional sketched linear regression framework. By analyzing multiplicative (signal-dependent) and additive (signal-independent) quantization, we identify a critical dichotomy in their scaling behaviors. Our analysis reveals that while both schemes introduce an additive error and degrade the effective data size, they exhibit distinct effects on effective model size: multiplicative quantization maintains the full-precision model size, whereas additive quantization reduces the effective model size. Numerical experiments validate our theoretical findings. By rigorously characterizing the complex interplay among model scale, dataset size, and quantization error, our work provides a principled theoretical basis for optimizing training protocols under practical hardware constraints.

低精度训练缩放定律线性回归量化理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。