揭示自回归模型预测下一个词背后的物理规律,连接信息与能量消耗。
Physics in Next-token Prediction
- 发现自回归模型中信息守恒定律,提出信息容量第一定律。
- 引入兰道尔原理,建立模型训练与能耗的第二定律关系。
- 理论统一了缩放定律、知识容量与精度缩放规律,适合模型效率研究者。
我们发现了自回归语言模型中下一个词预测(NTP)的底层物理规律。通过识别信息守恒定律,提出了信息容量第一定律(IC-1),表明自回归模型智能涌现的本质是信息传递过程。同时将兰道尔原理引入NTP,推导出信息容量第二定律(IC-2),建立了模型训练与能量消耗之间的关系。此外,我们推导出多个具有实际意义的推论,并验证了信息容量定律与神经语言模型缩放定律、知识容量缩放定律及精度缩放定律的一致性。
原文摘要 · Abstract (English)
We discovered the underlying physics in Next-token Prediction (NTP). We identified the law of information conservation within NTP and proposed the First Law of Information Capacity (IC-1), demonstrating that the essence of intelligence emergence in auto-regressive models is fundamentally a process of information transfer. We also introduced Landauer's Principle into NTP, formulating the Second Law of Information Capacity (IC-2), which establishes the relationship between auto-regressive model training and energy consumption. Additionally, we presented several corollaries, which hold practical significance for production practices. Finally, we demonstrate the consistency between the Law of Information Capacity and the Scaling Law for Neural Language Models, the Knowledge Capacity Scaling Laws, and the Scaling Laws for Precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。