研究小模型训练的计算瓶颈,帮资源有限团队省钱提效。
Computational Bottlenecks of Training Small-scale Large Language Models
- 测试不同超参数对20亿参数内模型训练的影响
- 发现通信协议和批处理大小显著影响每美元生成词数
- 适合预算紧张的科研机构优化训练配置
尽管大语言模型主导人工智能领域,但因成本与效率需求,小规模大语言模型(SLMs)正受到关注。然而,关于SLMs训练行为与计算需求的研究仍较有限。本研究通过分析多种超参数与配置(包括GPU类型、批处理大小、模型规模、通信协议、注意力机制及GPU数量)对训练过程的影响,探索了最多20亿参数的SLMs的计算瓶颈。我们在主流云服务上使用损失每美元和每秒生成词数等指标进行评估。研究结果旨在支持低资源人工智能研究机构更广泛地采用并优化语言模型训练。
原文摘要 · Abstract (English)
While large language models (LLMs) dominate the AI landscape, Small-scale large Language Models (SLMs) are gaining attention due to cost and efficiency demands from consumers. However, there is limited research on the training behavior and computational requirements of SLMs. In this study, we explore the computational bottlenecks of training SLMs (up to 2B parameters) by examining the effects of various hyperparameters and configurations, including GPU type, batch size, model size, communication protocol, attention type, and the number of GPUs. We assess these factors on popular cloud services using metrics such as loss per dollar and tokens per second. Our findings aim to support the broader adoption and optimization of language model training for low-resource AI research institutes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。