arXiv:2411.01663cs.LGcs.AI2024-11被引 5

首次证明1比特神经网络存在可扩展性定律,支持高效大模型发展

Unlocking the Theory Behind Scaling 1-Bit Neural Networks

  • 理论证明1比特网络随宽度增加趋于核方法行为
  • 模型宽度增大时损失趋近于零,泛化差距可忽略
  • 揭示训练损失与模型规模、数据量的幂律关系,适合硬件优化研究者

近期1比特大语言模型展现出媲美传统模型的性能与极致效率。研究表明,随着参数量增加,1比特模型性能持续提升,暗示可能存在1比特神经网络的缩放定律。本文首次严格证明该缩放定律:尽管权重被限制在{-1, +1},随着网络宽度增长,训练动态必然趋向核方法行为。这一理论突破保证了1比特模型在宽度增大时可收敛至任意小的损失。我们引入泛化差异概念,即1比特网络与全精度模型输出间的差距,证明该差距在宽度扩大时保持微小。基于Kaplan等(2020)的工作,进一步分析训练损失如何随模型规模、数据集大小及训练资源呈幂律变化。研究结果表明,1比特神经网络具备巨大扩展潜力,未来1比特(int1)或成标准精度。

原文摘要 · Abstract (English)

Recently, 1-bit Large Language Models (LLMs) have emerged, showcasing an impressive combination of efficiency and performance that rivals traditional LLMs. Research by Wang et al. (2023); Ma et al. (2024) indicates that the performance of these 1-bit LLMs progressively improves as the number of parameters increases, hinting at the potential existence of a Scaling Law for 1-bit Neural Networks. In this paper, we present the first theoretical result that rigorously establishes this scaling law for 1-bit models. We prove that, despite the constraint of weights restricted to $\{-1, +1\}$, the dynamics of model training inevitably align with kernel behavior as the network width grows. This theoretical breakthrough guarantees convergence of the 1-bit model to an arbitrarily small loss as width increases. Furthermore, we introduce the concept of the generalization difference, defined as the gap between the outputs of 1-bit networks and their full-precision counterparts, and demonstrate that this difference maintains a negligible level as network width scales. Building on the work of Kaplan et al. (2020), we conclude by examining how the training loss scales as a power-law function of the model size, dataset size, and computational resources utilized for training. Our findings underscore the promising potential of scaling 1-bit neural networks, suggesting that int1 could become the standard in future neural network precision.

1比特模型缩放定律理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。