arXiv:2503.01314cs.LGcs.AI2025-03被引 4

揭示大模型缩放定律在多元与核回归中的普遍性,深化对大模型学习机制的理解。

Scaling Law Phenomena Across Regression Paradigms: Multiple and Kernel Approaches

  • 拓展缩放定律至多元回归和核回归,突破线性模型限制。
  • 验证在多类回归中测试损失仍遵循幂律关系,跨越七数量级。
  • 为大模型泛化能力提供新解释,适合关注机器学习理论的读者。

近年来,大型语言模型(LLMs)取得了显著成功,其背后的关键因素是OpenAI发现的缩放定律。具体而言,对于Transformer架构的模型,测试损失与模型规模、数据集规模及训练计算量之间呈现幂律关系,覆盖超过七个数量级的变化范围。这一现象挑战了传统机器学习的奥司卡剪刀原理,即过参数化算法会在训练集上过拟合,导致测试性能下降。尽管已有研究在简单机器学习场景(如线性回归)中发现了缩放定律,但对实际大模型中该现象的完整解释仍不明确。本文进一步推进理解,证明缩放定律现象同样存在于更强大且表达力更强的多元回归与核回归设置中。我们的分析为缩放定律提供了更深层洞见,有望增进对大型语言模型学习行为的认识。

原文摘要 · Abstract (English)

Recently, Large Language Models (LLMs) have achieved remarkable success. A key factor behind this success is the scaling law observed by OpenAI. Specifically, for models with Transformer architecture, the test loss exhibits a power-law relationship with model size, dataset size, and the amount of computation used in training, demonstrating trends that span more than seven orders of magnitude. This scaling law challenges traditional machine learning wisdom, notably the Oscar Scissors principle, which suggests that an overparametrized algorithm will overfit the training datasets, resulting in poor test performance. Recent research has also identified the scaling law in simpler machine learning contexts, such as linear regression. However, fully explaining the scaling law in large practical models remains an elusive goal. In this work, we advance our understanding by demonstrating that the scaling law phenomenon extends to multiple regression and kernel regression settings, which are significantly more expressive and powerful than linear methods. Our analysis provides deeper insights into the scaling law, potentially enhancing our understanding of LLMs.

缩放定律多元回归核方法大模型理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。