arXiv:2510.04067cs.LGcs.AI2025-10被引 1

发现交叉熵下降变慢的根源在于其隐藏成分的变化差异。

What Scales in Cross-Entropy Scaling Law?

  • 将交叉熵分解为误差熵、自对齐和置信度三部分,揭示真实变化机制。
  • 仅误差熵遵循幂律缩放,其他两项基本不变,且占比随模型增大而下降。
  • 提出误差熵缩放定律,更适合指导大模型训练与未来发展。

交叉熵缩放定律长期用于指导大语言模型研发,表明损失随模型规模增大呈可预测的幂律下降。但近期研究表明,在超大规模下该定律失效:损失下降速度远低于预期。本文假设根本原因在于交叉熵本身并不真正缩放,而是其隐藏组件之一在变化。为此,我们提出一种新的交叉熵分解方法,将其拆分为误差熵、自对齐和置信度三部分。理论与实证均表明该分解能精确捕捉训练动态与优化目标。在多个数据集及32个跨越五数量级规模的模型上进行广泛实验发现,仅误差熵呈现稳健的幂律缩放,其余两项基本保持不变。且误差熵在小模型中占交叉熵主导地位,随模型增大比例逐渐降低。这解释了为何交叉熵缩放定律在小规模有效而大规模失效。我们的研究确立了误差熵缩放定律作为更准确的模型行为描述,有望在大模型训练、理解与未来发展中广泛应用。

原文摘要 · Abstract (English)

The cross-entropy scaling law has long served as a key tool for guiding the development of large language models. It shows that cross-entropy loss decreases in a predictable power-law rate as the model size increases. However, recent evidence indicates that this law breaks down at very large scales: the loss decreases more slowly than expected, which causes significant trouble for developing large language models. In this paper, we hypothesize that the root cause lies in the fact that cross-entropy itself does not truly scale; instead, only one of its hidden components does. To investigate this, we introduce a novel decomposition of cross-entropy into three parts: Error-Entropy, Self-Alignment, and Confidence. We show both theoretically and empirically that this decomposition precisely captures the training dynamics and optimization objectives. Through extensive experiments on multiple datasets and 32 models spanning five orders of magnitude in size, we find that only error-entropy follows a robust power-law scaling, while the other two terms remain largely invariant. Moreover, error-entropy constitutes the dominant share of cross-entropy in small models but diminishes in proportion as models grow larger. This explains why the cross-entropy scaling law appears accurate at small scales but fails at very large ones. Our findings establish the error-entropy scaling law as a more accurate description of model behavior. We believe it will have wide applications in the training, understanding, and future development of large language models.

大模型缩放定律交叉熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。