用逆问题方法高效发现大模型缩放规律,降低训练成本。
Uncovering Scaling Laws for Large Language Models via Inverse Problems
- 将逆问题思想引入大模型研究,反推缩放规律。
- 揭示数据量与计算量的最优配比关系,提升效率。
- 适合关注模型训练优化与资源节约的研究者。
大型语言模型(LLMs)在多个领域取得显著成功,其背后是数据和计算规模的空前增长。然而,由于训练成本高昂,传统的试错法难以有效改进模型。受逆问题在科学定律发现中成功应用的启发,本文提出通过逆问题方法高效揭示指导大模型构建的缩放规律,从而在显著提升性能的同时实现更高的成本效益。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are large-scale pretrained models that have achieved remarkable success across diverse domains. These successes have been driven by unprecedented complexity and scale in both data and computations. However, due to the high costs of training such models, brute-force trial-and-error approaches to improve LLMs are not feasible. Inspired by the success of inverse problems in uncovering fundamental scientific laws, this position paper advocates that inverse problems can also efficiently uncover scaling laws that guide the building of LLMs to achieve the desirable performance with significantly better cost-effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。