首次证明1比特神经网络存在可扩展性定律,支持高效大模型发展
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
- 理论证明1比特网络随宽度增加趋于核方法行为
- 模型宽度增大时损失趋近于零,泛化差距可忽略
- 揭示训练损失与模型规模、数据量的幂律关系,适合硬件优化研究者
近期1比特大语言模型展现出媲美传统模型的性能与极致效率。研究表明,随着参数量增加,1比特模型性能持续提升,暗示可能存在1比特神经网络的缩放定律。本文首次严格证明该缩放定律:尽管权重被限制在{-1, +1},随着网络宽度增长,训练动态必然趋向核方法行为。这一理论突破保证了1比特模型在宽度增大时可收敛至任意小的损失。我们引入泛化差异概念,即1比特网络与全精度模型输出间的差距,证明该差距在宽度扩大时保持微小。基于Kaplan等(2020)的工作,进一步分析训练损失如何随模型规模、数据集大小及训练资源呈幂律变化。研究结果表明,1比特神经网络具备巨大扩展潜力,未来1比特(int1)或成标准精度。
原文摘要 · Abstract (English)
Recently, 1-bit Large Language Models (LLMs) have emerged, showcasing an impressive combination of efficiency and performance that rivals traditional LLMs. Research by Wang et al. (2023); Ma et al. (2024) indicates that the performance of these 1-bit LLMs progressively improves as the number of parameters increases, hinting at the potential existence of a Scaling Law for 1-bit Neural Networks. In this paper, we present the first theoretical result that rigorously establishes this scaling law for 1-bit models. We prove that, despite the constraint of weights restricted to $\{-1, +1\}$, the dynamics of model training inevitably align with kernel behavior as the network width grows. This theoretical breakthrough guarantees convergence of the 1-bit model to an arbitrarily small loss as width increases. Furthermore, we introduce the concept of the generalization difference, defined as the gap between the outputs of 1-bit networks and their full-precision counterparts, and demonstrate that this difference maintains a negligible level as network width scales. Building on the work of Kaplan et al. (2020), we conclude by examining how the training loss scales as a power-law function of the model size, dataset size, and computational resources utilized for training. Our findings underscore the promising potential of scaling 1-bit neural networks, suggesting that int1 could become the standard in future neural network precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。