arXiv:2504.15208cs.LGcs.AI2025-04ICLR被引 7

更大语言模型泛化更好,因计算最优下误差随规模下降。

Compute-Optimal LLMs Provably Generalize Better With Scale

  • 基于计算最优规律,构建损失方差与量化误差的泛化界。
  • 规模增大时,每样本参数不变,但损失方差和量化误差下降。
  • 揭示大模型更易量化,适合研究模型缩放与泛化关系者阅读。

为探究为何大语言模型泛化能力更强,本文在符合Chinchilla缩放定律的计算最优范式下,推导了大语言模型预训练目标的泛化界。提出一种全新的全经验弗里德曼型鞅集中不等式,通过考虑损失函数方差,收紧了现有边界。该泛化界可分解为三部分:每标记的参数量、损失方差和固定比特率下的量化误差。在计算最优路径上,随着模型规模扩大,每数据点的参数量保持恒定;但损失方差与量化误差均下降,表明大模型应具有更小的泛化差距。从信息论角度分析发现,大模型整合新信息的速率增长慢于其容量增长。据此建立了泛化差距的缩放律,边界随规模扩大而更可预测地增强。

原文摘要 · Abstract (English)

Why do larger language models generalize better? To investigate this question, we develop generalization bounds on the pretraining objective of large language models (LLMs) in the compute-optimal regime, as described by the Chinchilla scaling laws. We introduce a novel, fully empirical Freedman-type martingale concentration inequality that tightens existing bounds by accounting for the variance of the loss function. This generalization bound can be decomposed into three interpretable components: the number of parameters per token, the loss variance, and the quantization error at a fixed bitrate. As compute-optimal language models are scaled up, the number of parameters per data point remains constant; however, both the loss variance and the quantization error decrease, implying that larger models should have smaller generalization gaps. We examine why larger models tend to be more quantizable from an information theoretic perspective, showing that the rate at which they can integrate new information grows more slowly than their capacity on the compute-optimal frontier. From these findings we produce a scaling law for the generalization gap, with bounds that become predictably stronger with scale.

大模型泛化能力缩放定律量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。