arXiv:2602.07238cs.AIcs.LG2026-02被引 3

大模型性能提升主要靠算力,但公司间效率差异显著。

Is there "Secret Sauce'' in Large Language Model Development?

  • 通过809个模型数据建模,分离出开发者特有效率优势。
  • 前沿模型性能90%由算力决定,非技术秘密。
  • 公司内部模型效率差超40倍,适合关注训练优化的团队。

基于2022至2025年间发布的809个大语言模型的训练与基准数据,我们采用包含发布时间和开发方固定效应的缩放定律回归分析。结果表明,开发者存在显著的效率优势,但其重要性取决于模型所处性能分布位置。在模型性能前沿,80%-90%的性能差异由更高训练算力解释,说明前沿进步主要依赖规模而非专有技术。而在远离前沿时,专有技术和共享算法进步可大幅降低达到特定能力阈值所需的算力。部分公司能系统性地更高效地训练更小模型。尤为惊人的是,同一公司内部模型间效率差异可达40倍以上。研究还讨论了对人工智能领导地位与能力扩散的影响。

原文摘要 · Abstract (English)

Do leading LLM developers possess a proprietary ``secret sauce'', or is LLM performance driven by scaling up compute? Using training and benchmark data for 809 models released between 2022 and 2025, we estimate scaling-law regressions with release-date and developer fixed effects. We find clear evidence of developer-specific efficiency advantages, but their importance depends on where models lie in the performance distribution. At the frontier, 80-90% of performance differences are explained by higher training compute, implying that scale--not proprietary technology--drives frontier advances. Away from the frontier, however, proprietary techniques and shared algorithmic progress substantially reduce the compute required to reach fixed capability thresholds. Some companies can systematically produce smaller models more efficiently. Strikingly, we also find substantial variation of model efficiency within companies; a firm can train two models with more than 40x compute efficiency difference. We also discuss the implications for AI leadership and capability diffusion.

大模型算力效率模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。