arXiv:2509.23963cs.LG2025-09被引 3

验证了Chinchilla缩放原则的稳健性,证明其核心结论不受参数定义影响。

Evaluating the Robustness of Chinchilla Compute-Optimal Scaling

  • 重新分析模型参数定义,发现三种解释差异达15.2%
  • 不同参数定义下,最优训练量与参数比仍保持稳定
  • 即使大幅扰动参数,关键结论依然成立,适合模型缩放决策

Hoffman等(2022)提出的Chinchilla计算最优缩放原则为语言模型扩展奠定了基础。然而,近年来对其有效性提出质疑:置信区间过宽、三种方法间存在分歧,且与其他缩放规律不一致。这引发关键问题:从业者是否仍可依赖Chinchilla的指导?本研究表明答案是肯定的。我们发现,支撑Chinchilla分析的核心模型参数存在歧义:三种可能解释下参数差异最高达15.2%。但出人意料的是,不同参数解释对关键结果——缩放律估计和计算最优的训练样本-参数比——影响甚微;在一种解释下,该比例甚至随目标计算预算更趋恒定。进一步通过四种结构化方式人为扰动参数,发现关键结果最敏感于加性或系统性误差,这类误差会破坏最优比例的平坦趋势,但整体而言,主要结论仍能承受显著扰动。综上,本研究为Chinchilla作为语言模型扩展的可靠指南提供了新的信心。

原文摘要 · Abstract (English)

Hoffman et al (2022)'s Chinchilla paper introduced the principle of compute-optimal scaling, laying a foundation for future scaling of language models. In the years since, however, valid concerns about Chinchilla have been raised: wide confidence intervals, discrepancies between its three approaches, and incongruities with other scaling laws. This raises a critical question for the field: Can practitioners still rely on Chinchilla's prescriptions? Our work demonstrates the answer is yes. We begin by uncovering that the model parameters central to Chinchilla's analyses were ambiguous: three interpretations are possible, with relative differences between different interpretations of model parameters as high as 15.2%. We find that, perhaps surprisingly, which model parameters are used for the analyses do not meaningfully affect key results: the scaling law estimates and the compute-optimal tokens-to-parameter ratio. Indeed, under one interpretation, the tokens-to-parameter ratio becomes more constant with the target compute budget. We then ask how distorted the Chinchilla model parameters could have been without meaningfully affecting the key results. By deliberately perturbing model parameters in four structured ways, we find that key Chinchilla results are most sensitive to additive or systematic errors, which can alter the otherwise flat trend of the optimal tokens-to-parameter ratio, but overall, Chinchilla's key results withstand sizable perturbations. Altogether, our findings offer the field renewed confidence in Chinchilla as a durable guide for scaling language models.

模型缩放计算最优鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。