arXiv:2606.25008cs.LGcs.CL2026-06

固定指数下,系数决定大模型性能上限。

Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients

论文配图:Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
图 1 · 摘自论文原文
  • 用通用机制解释神经网络缩放律的指数为何固定。
  • 系数受数据与架构影响,直接决定最优模型形状。
  • 关注系数是提升实际性能的关键突破口。

神经网络缩放律描述了预训练损失随训练时间、模型规模和计算量呈幂律下降。本文认为这些幂律的指数由通用机制决定:由于Softmax强非线性导致时间缩放指数为1/3,由于表示叠加导致宽度反比缩放,由于Transformer层的集成平均导致深度反比缩放。这些机制对多种数据结构和架构细节具有鲁棒性,使当前大语言模型处于具有固定指数的普适性类中。然而,系数对数据和架构细节敏感,直接影响最优模型形状和计算最优边界。因此,理解系数是实现近期性能提升的关键,进一步探索当前普适性类或可发现更优的普适性类。

原文摘要 · Abstract (English)

Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute. This position paper argues that the exponents of these power laws are fixed by generic mechanisms: a one-third time scaling due to the strong nonlinearity of Softmax, an inverse width scaling due to representational superposition, and an inverse depth scaling due to ensemble averaging of Transformer layers. These mechanisms are robust to a wide range of data structures and architectural details, placing current large language models in a universality class with fixed exponents. The coefficients, however, are expected to be sensitive to data and architecture details, and directly determine practical quantities such as the optimal model shape and the compute-optimal frontier. We therefore argue that understanding the coefficients is the key to near-term performance improvements, and that a closer examination of the current universality class may reveal pathways to better universality classes.

缩放律大模型普适性系数分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。