揭示大模型缩放指数过小背后的能源不可持续性问题
On the Smallness of the Large Language Models Scaling Exponents
- 提出缩放指数过小源于数据无限时损失函数的非零截距('基底效应')
- 证明即使修正该偏差,能源可持续性问题依然存在
- 类比流体湍流粗糙度,解释数据平滑性对缩放行为的影响
我们讨论当前大语言模型应用中缩放指数偏小的原因,指出其在能源资源方面处于不可持续状态。进一步表明,将该现象归因于无限数据极限下损失函数非零值('基底效应')所导致的数值偏差,并不能消除不可持续性问题。最后,基于流体湍流的拟经验模型类比,讨论了数据平滑性(粗糙性)对缩放指数的影响。
原文摘要 · Abstract (English)
We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are indicating an unsustainable regime in terms of energy resources. We further show that attributing the smallness of such exponents to a numerical bias due to the neglect of a non-zero value of the loss function in the limit of infinite data (``pedestal effect") does not remove the unsustainability issue. Finally, the effects of the smoothness (roughness) of the data on the scaling exponents is commented upon based on an analogy with phenomenological models of fluid turbulence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。