arXiv:2412.16443cs.LGcs.AI2024-12被引 4

揭示大模型缩放规律,指出当前面临回报递减与资源瓶颈。

Has LLM Reached the Scaling Ceiling Yet? Unified Insights into LLM Regularities and Constraints

  • 用中心极限定理解释隐藏层噪声随上下文变长而降低
  • 发现预测误差由熵、偏差和方差构成,缩放效果逐渐减弱
  • 提出信噪比阈值,说明能力突现需达到一定质量门槛

大型语言模型展现出惊人能力,但其可扩展性引发关键问题:是否已达缩放极限?本文构建统一理论框架,融合数学与统计洞察,解释大模型缩放动态。首先,提出隐藏表示的中心极限定理,表明隐藏层噪声随上下文长度增加而反比下降,解释了上下文长度优化的稳定效应与局限;其次,通过偏差-方差分解,将下一词预测误差拆分为不可约熵、容量驱动偏差与有限样本方差,揭示缩放带来的边际收益递减;第三,定义信噪比(SNR),量化能力在SNR突破阈值时的突现现象,指出缩放有效性随质量不足而下降。综合表明,尽管大模型尚未触及绝对缩放天花板,但实际约束日益显著:回报递减、资源低效与数据限制。未来进展需从盲目缩放转向架构创新、数据质量提升与训练范式革新。本工作为下一代大模型高效发展提供路线图,推动领域超越传统缩放策略。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their scalability raises a critical question: Have we reached the scaling ceiling? This paper addresses this pivotal question by developing a unified theoretical framework that integrates mathematical and statistical insights to explain the scaling dynamics of LLMs. We present: 1. Central Limit Theorem (CLT) for Hidden Representations: We show that noise in hidden representations scales inversely with context size, explaining stabilization effects and the limits of context length improvements. 2. Bias-Variance Decomposition: We decompose next-token prediction loss into irreducible entropy, capacity-driven bias, and finite sample variance, revealing trade-offs where scaling yields diminishing returns. 3. Emergent SNR Thresholds: By defining signal-to-noise ratio (SNR), we quantify how capabilities emerge abruptly once SNR surpasses a threshold, offering insights into when scaling becomes less effective. Through this framework, we conclude that while LLMs have not reached an absolute scaling ceiling, practical constraints are increasingly prominent: diminishing returns, resource inefficiencies, and data limitations. Future progress will require a shift from brute-force scaling to innovations in architecture, data quality, and training paradigms. This work provides a roadmap for guiding the efficient development of next-generation LLMs and advancing the field beyond traditional scaling strategies. Keywords: Large Language Models; Scaling Ceiling; Central Limit Theorem; Bias-Variance Trade-Off; Signal-to-Noise Ratio; Emergent Capabilities

大模型缩放规律信噪比偏差-方差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。