arXiv:2505.10465cs.LGcs.AI2025-05NeurIPS被引 29

发现大模型性能随规模提升的关键是特征超叠加现象。

Superposition Yields Robust Neural Scaling

  • 用权重衰减控制特征超叠加程度,系统研究模型规模与损失关系。
  • 强超叠加下损失与模型维度成反比,适用多种数据分布。
  • 实证开源大模型符合该规律,解释了经典缩放定律的成因。

当前大语言模型的成功依赖于模型越大性能越好的现象。然而,损失随模型规模呈幂律下降的神经缩放规律来源尚不明确。我们提出,表示超叠加(即模型表示的特征数超过其维度)可能是导致损失下降和神经缩放的关键因素。基于 Anthropic 的玩具模型,我们通过权重衰减控制超叠加程度,系统研究损失随模型规模的变化。当超叠加较弱时,仅当数据特征频率呈幂律分布,损失才遵循幂律;而在强超叠加下,损失在广泛频率分布下均与模型维度成反比,这是由于表示向量间的几何重叠所致。我们验证了开源大模型处于强超叠加状态,其损失缩放与模型维度成反比,且 Chinchilla 缩放定律也与此行为一致。结果表明,表示超叠加是神经缩放规律的核心驱动因素,为理解缩放规律何时可改进或失效提供了洞见。

原文摘要 · Abstract (English)

The success of today's large language models (LLMs) depends on the observation that larger models perform better. However, the origin of this neural scaling law, that loss decreases as a power law with model size, remains unclear. We propose that representation superposition, meaning that LLMs represent more features than they have dimensions, can be a key contributor to loss and cause neural scaling. Based on Anthropic's toy model, we use weight decay to control the degree of superposition, allowing us to systematically study how loss scales with model size. When superposition is weak, the loss follows a power law only if data feature frequencies are power-law distributed. In contrast, under strong superposition, the loss generically scales inversely with model dimension across a broad class of frequency distributions, due to geometric overlaps between representation vectors. We confirmed that open-sourced LLMs operate in the strong superposition regime and have loss scaling inversely with model dimension, and that the Chinchilla scaling laws are also consistent with this behavior. Our results identify representation superposition as a central driver of neural scaling laws, providing insights into questions like when neural scaling laws can be improved and when they will break down.

大模型缩放定律超叠加模型性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。